BuildNot yet confirmed elsewhere1 publisher2 min readPublished
Before you restart that Windows service, find out whether StartType survived the update
A five-move walkthrough of a stopped SQL Server argues for evidence before action. The restart that clears the alert also rotates the log that would have explained the repeat.
The Engineer · Build desk
What happened
- A dev.to walkthrough reduces any stopped Windows service to the same five investigative moves, and puts "do not restart first" ahead of all of them.
- Move one is a Get-Service read of Status and StartType, on the grounds that OS updates sometimes knock StartType off Automatic.
- Event ID 7036 in the System log supplied the exact stop and restart times, which the author sums as a six-minute outage matching the alert.
- Event 1074 showed two reboots in one maintenance window: one triggered by an Ansible account, then a TrustedInstaller reboot marked "Operating System: Upgrade (Planned)".
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- cost The author's price for the evidence is five minutes of read-only queries. The alternative is paying for the same outage again whenever the automation that caused it next runs.
- decision A runbook that starts with a restart converts every alert into an incident and every incident into the same fix, which is what makes recurrence look like bad luck rather than a setting.
- exposure When a start setting is left wrong, the people who control the patch and configuration automation decide when it becomes an outage, not the team that owns the service.
Restarting is not a read-only act, and on the host in this case it removes the two artifacts that explain a repeat. SQL Server rotates its ERRORLOG on every restart [13], so a bounce pushes the outage window out of the current file and turns the instruction to read the entries immediately before the stop [12] into a search by LastWriteTime [13]. Event 7036 records every service state change [8], including the one the responder just caused, and the pattern the guide teaches is several stop/start pairs in quick succession as the signature of a crash-and-restart loop [9]. A responder's stop followed by a start has the same shape.
StartType is where the recurrence hides. It shares the first query with Status because operating system updates occasionally reset it away from Automatic [6]. A manual restart returns the service to Running and leaves that field exactly as the update left it, so the alert clears and returns at the next reboot with nothing unusual in the application log to point at [2]. The author's read of the SQL case is that Windows Update had quietly queued a second reboot behind the Ansible one, and that without Event 1074 the outage would have looked unexplained [1]. Put the two findings together and you have the failure the restart cannot reach: the update changes the start setting, then supplies the reboot that tests it [17].
The reason the rule is cheap to follow sits in the same case. All three SQL services came back Running and Automatic when the first query ran, so the alert was valid but there was nothing left to act on [7], and the monitoring poll interval is why it outlived the outage [5]. A restart on arrival would have taken a running service down and written that stop into the log everyone was about to read [8].
One number worth checking before it reaches a customer note: the transitions are logged at 10:20 AM IST and 10:27 AM IST [10], which is seven minutes, and the write-up calls it six [15][16]. Durations belong to TimeCreated, not to the prose wrapped around it. The published walkthrough also breaks off mid-command at the ERRORLOG step [14], so the case never shows which of the three application-log signatures the SQL instance actually carried [2], and that is the part that would have separated a crash from a host that took the process down with it.
What to watch
- Whether StartType on MSSQLSERVER, SQLSERVERAGENT and MsDtsServer130 is still Automatic after the next cumulative update on that host.
- Whether the next patch window again puts a TrustedInstaller reboot behind the Ansible one, which would make the double reboot a scheduling defect rather than a one-off.
- A finished version of the ERRORLOG step, showing whether the 10:20 stop was a clean shutdown or an exception.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence34
- Adoption15
- Hype gap+14
- Incentives24
- Confidence42
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
The author concludes that Windows Update had quietly queued a second reboot on top of the Ansible reboot, and that without checking Event 1074 this would have looked like an unexplained outage.
ReportedSupportedSource: dev.to walkthrough2 sources— create a free account to open themView cited source - [2]
The application log carries one of three signatures: an error or exception before the stop means a crash, a clean shutdown message means a planned stop, and nothing unusual means an external trigger such as a reboot killing the process.
- [3]
The walkthrough presents the investigation of a stopped Windows service as the same five moves whether the service is SQL Server, IIS, a background agent or a custom app.
- [4]
The article's first rule is not to restart first: restarting without knowing why the service stopped can hide a failing update, an exhausted host, or an automation script that will stop it again in the next cycle. It tells the reader to spend five minutes on evidence first.
- [5]
The first question is whether the service is still stopped, because monitoring systems have polling intervals and the service may have recovered before the alert arrived; ground truth comes before any action.
- [6]
The first check is a Get-Service query returning Name, Status and StartType. Status shows whether the service is down now; StartType matters because OS updates occasionally reset it from Automatic.
- [7]
In the SQL Server example, MSSQLSERVER, SQLSERVERAGENT and MsDtsServer130 all returned Running / Automatic, having self-recovered before anyone looked, which told the team the alert was valid but not actionable.
- [8]
Windows logs every service state change (stopped, running, paused) as Event ID 7036 in the System log; the guide pulls the last 20 transitions and sorts them chronologically to get exact timestamps.
- [9]
Multiple stop/start cycles in quick succession in the 7036 output often mean a crash-and-restart loop.
- [10]
The 7036 output showed all three SQL services stopped at 10:20 AM IST and back to running at 10:27 AM IST.
- [11]
Event 1074 revealed two reboots in the same maintenance window: the first triggered by o9ansibleuser via Ansible, the second chained by TrustedInstaller.exe with reason "Operating System: Upgrade (Planned)".
- [12]
Windows event logs tell you when the service stopped; the application's own log tells you why, and the guide directs the reader to the entries immediately before the stop.
- [13]
SQL Server rotates its log on every restart, so the file covering an outage window has to be found by LastWriteTime.
- [14]
The published walkthrough breaks off mid-command in the SQL Server ERRORLOG file-discovery example.
- [15]
The article describes that sequence as a 6-minute outage that overlapped exactly with the alert timestamp.
- [16]
The interval between the two logged timestamps is seven minutes, one minute more than the six-minute outage the article reports.
- [17]
An OS update that resets StartType away from Automatic and also queues a host reboot produces a service that does not come back, with no application-log error behind it.
Sources
1 independent publisher whose own reporting we read for this story.
- dev.toA Windows Service is Down. Now What?
1 article · August 23, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
Entities
- Microsoft SQL Server (MSSQL)Follow
- PowerShellFollow
- AnsibleFollow
- Windows UpdateFollow
- Windows Service Control ManagerFollow
- Windows Event ID 7036Follow
- Windows Event ID 1074Follow
- Windows Event ID 6008Follow
- Internet Information ServicesFollow
- dev.toFollow