sb-tailscale stuck in restart loop — clean logins each cycle, healthcheck timing suspected
Posted by DemonsElemental
Hello SparkBox team! First off, I want to say I appreciate everything you are doing with this project. I have been looking for a solution like this for years, and your quality support is the cherry on the cake! I cant overstate how nice it is to have people willing to take the time to explain the process and work with people like me who are trying to self-teach. I have been able to fix all my problems to date using direction TomAI, but this one the AI couldn't find a solution for. Please see the summery and logs below. Symptom: sb-tailscale container is continuously restarting (currently seen at "Up 6s", health: starting → then bounced again). Not a one-off — confirmed it cycling live. What the logs show: Each restart, tailscaled comes up cleanly — authenticates, logs in, connects out to DERP with no errors. No bad auth key, no config error, no network/firewall block visible in the log output. It never logs a crash or fatal error before the container gets killed and restarted. What this looks like: Because every cycle ends in a clean, successful connection rather than a crash, this points to the container's healthcheck window being too short for how long tailscaled actually takes to mark itself ready — something (likely SparkBox's deunhealth auto-heal watcher) is killing/restarting it before the healthcheck has a chance to pass, creating a self-inflicted loop rather than a real fault. Impact: Cosmetic-ish / annoying rather than broken — each cycle it does successfully authenticate and connect to the tailnet before getting bounced again, so connectivity isn't fully dead, just unstable. Tried: Did not blind-restart the container since logs don't show a wedged/hung state — restarting looked unlikely to hold past the next healthcheck cycle given the pattern. Ask: Healthcheck interval/threshold for the tailscale container may need tuning so it doesn't get killed mid-startup. 2026/08/01 00:29:59 localapi: [POST] /localapi/v0/debug 2026/08/01 00:30:14 localapi: [POST] /localapi/v0/debug boot: 2026/08/01 00:30:28 Sending SIGTERM to tailscaled 2026/08/01 00:30:28 tailscaled got signal terminated; shutting down 2026/08/01 00:30:28 canceling captive portal context 2026/08/01 00:30:28 control: client.Shutdown ... 2026/08/01 00:30:28 control: updateRoutine: exiting 2026/08/01 00:30:28 control: authRoutine: exiting 2026/08/01 00:30:28 control: mapRoutine: exiting 2026/08/01 00:30:28 control: Client.Shutdown done. 2026/08/01 00:30:28 ipnext: work queue shutdown failed: execqueue shut down 2026/08/01 00:30:28 magicsock: closing connection to derp-9 (conn-close), age 1m30s 2026/08/01 00:30:28 magicsock: 0 active derp conns 2026/08/01 00:30:28 flushing log. 2026/08/01 00:30:28 logger closing down boot: 2026/08/01 00:30:28 tailscaled exited boot: 2026/08/01 00:30:29 Starting tailscaled boot: 2026/08/01 00:30:29 Waiting for tailscaled socket at /tmp/tailscaled.sock TPM: error opening: stat /dev/tpmrm0: no such file or directory 2026/08/01 00:30:29 logtail started 2026/08/01 00:30:29 Program starting: v1.98.10-t36550d57f, Go 1.26.5: []string{"tailscaled", "--socket=/tmp/tailscaled.sock", "--statedir=/var/lib/tailscale", "--tun=userspace-networking"} 2026/08/01 00:30:29 LogID: cff5ae0f782e785a74cdd9809e64dbb25dbbadd72807026fae4d2a6fb66960bc 2026/08/01 00:30:29 logpolicy: using system state directory "/var/lib/tailscale" 2026/08/01 00:30:29 dns: [rc=unknown ret=direct] 2026/08/01 00:30:29 dns: using "direct" mode 2026/08/01 00:30:29 dns: using dns.directManager 2026/08/01 00:30:29 dns: inotify: NewDirWatcher: context canceled 2026/08/01 00:30:29 wgengine.NewUserspaceEngine(tun "userspace-networking") ... 2026/08/01 00:30:29 tstun: error initializing tun dev stats polling: error getting ifIndex: no such device 2026/08/01 00:30:29 dns: using dns.noopManager 2026/08/01 00:30:29 Creating WireGuard device... 2026/08/01 00:30:29 Bringing WireGuard device up... 2026/08/01 00:30:29 Bringing router up... 2026/08/01 00:30:29 Clearing router settings... 2026/08/01 00:30:29 Starting network monitor... 2026/08/01 00:30:29 Engine created. 2026/08/01 00:30:29 pm: using backend prefs for "profile-7d9d": Prefs{ra=true dns=false want=true routes=[0.0.0.0/0 ::/0] snat=true statefulFiltering=false nf=on update=check Persist{o=, n=[UJAAj] u="drace999@gmail.com" ak=-}} 2026/08/01 00:30:29 logpolicy: using system state directory "/var/lib/tailscale" 2026/08/01 00:30:29 linkChange: in state NoState; PAC or proxyConfig changed; updating routes 2026/08/01 00:30:29 got LocalBackend in 14ms 2026/08/01 00:30:29 Start 2026/08/01 00:30:29 ipnext: "conn25": skipping extension 2026/08/01 00:30:29 ipnext: active extensions: conn25, portlist, posture, clientupdate, relayserver, taildrop 2026/08/01 00:30:29 load netmap from cache: netmap cache is not available 2026/08/01 00:30:29 Backend: logs: be:cff5ae0f782e785a74cdd9809e64dbb25dbbadd72807026fae4d2a6fb66960bc fe: 2026/08/01 00:30:29 control: client.Login(0) 2026/08/01 00:30:29 control: doLogin(regen=false, hasUrl=false) 2026/08/01 00:30:29 health(warnable=warming-up): error: Tailscale is starting. Please wait. boot: 2026/08/01 00:30:29 [warning] failed to symlink socket: file exists To interact with the Tailscale CLI please use tailscale --socket="/tmp/tailscaled.sock" boot: 2026/08/01 00:30:29 tailscaled in state "NoState", waiting 2026/08/01 00:30:29 control: control server key from https://controlplane.tailscale.com: ts2021=[fSeS+], legacy=[nlFWp] 2026/08/01 00:30:29 control: RegisterReq: onode= node=[UJAAj] fup=false nks=false 2026/08/01 00:30:29 control: RegisterReq: got response; nodeKeyExpired=false, machineAuthorized=true; authURL=false 2026/08/01 00:30:29 health(warnable=not-in-map-poll): ok 2026/08/01 00:30:29 control: netmap: got new dial plan from control 2026/08/01 00:30:29 netmap: suggested exit node: no preferred DERP, try again later 2026/08/01 00:30:29 Switching ipn state NoState - Starting (WantRunning=true, nm=true) 2026/08/01 00:30:29 magicsock: SetPrivateKey called (init) 2026/08/01 00:30:29 wgengine: Reconfig: configuring userspace WireGuard config (with 1 peers) 2026/08/01 00:30:29 wgengine: Reconfig: configuring router 2026/08/01 00:30:29 wgengine: Reconfig: user dialer 2026/08/01 00:30:29 tsdial: bart table size: 4 2026/08/01 00:30:29 wgengine: Reconfig: configuring DNS 2026/08/01 00:30:29 dns: Set: {DefaultResolvers:[] Routes:{} SearchDomains:[] Hosts:2} 2026/08/01 00:30:29 dns: Resolvercfg: {Routes:{} Hosts:2 LocalDomains:[]} 2026/08/01 00:30:29 dns: OScfg: {} boot: 2026/08/01 00:30:29 tailscaled in state "Starting", waiting 2026/08/01 00:30:30 magicsock: home DERP changing from derp-0 [0ms] to derp-13 [13ms] (forced=false) 2026/08/01 00:30:30 magicsock: home is now derp-13 (den) 2026/08/01 00:30:30 magicsock: adding connection to derp-13 for home-keep-alive 2026/08/01 00:30:30 magicsock: 1 active derp conns: derp-13=cr0s,wr0s 2026/08/01 00:30:30 derphttp.Client.Connect: connecting to derp-13 (den) 2026/08/01 00:30:30 control: NetInfo: NetInfo{varies=false ipv6=false ipv6os=true udp=true icmpv4=false derp=13 portmap=active-UMC link="" firewallmode=""} 2026/08/01 00:30:30 writing netmap to disk cache 2026/08/01 00:30:30 health(warnable=no-derp-connection): ok 2026/08/01 00:30:30 Switching ipn state Starting - Running (WantRunning=true, nm=true) 2026/08/01 00:30:30 health(warnable=warming-up): ok 2026/08/01 00:30:30 health(warnable=no-derp-connection): ok boot: 2026/08/01 00:30:30 Running 'tailscale set' 2026/08/01 00:30:30 localapi: [POST] /localapi/v0/check-prefs 2026/08/01 00:30:30 localapi: [PATCH] /localapi/v0/prefs boot: 2026/08/01 00:30:30 serve proxy: unsetting previous config 2026/08/01 00:30:30 localapi: [POST] /localapi/v0/serve-config 2026/08/01 00:30:30 health(warnable=no-derp-connection): ok 2026/08/01 00:30:30 [RATELIMIT] format("health(warnable=%s): ok") 2026/08/01 00:30:30 magicsock: derp-13 connected; connGen=1 2026/08/01 00:30:45 localapi: [POST] /localapi/v0/debug boot: 2026/08/01 00:30:45 Startup complete, waiting for shutdown signal boot: 2026/08/01 00:30:45 serve proxy: this node is configured as a proxy that exposes an HTTPS endpoint to tailnet, (perhaps a Kubernetes operator Ingress proxy) but it is not able to issue TLS certs, so this will likely not work. To make it work, ensure that HTTPS is enabled for your tailnet, see https://tailscale.com/kb/1153/enabling-https for more details.
4 replies
Chris wrote:
Nice diagnostic work here, and good news first: your login and connection are totally clean, nothing wrong with your setup. I checked the restart-timing theory against our actual healthcheck settings though, and the math doesn't line up: your own logs show Tailscale fully connected around 16 seconds in, well inside the window we give it, so that's probably not what's killing it. My next guess is something forcing the process to stop rather than it crashing on its own, maybe it's running low on the memory we allow it. Right after the next restart, run this and paste what it prints: sudo docker inspect sb-tailscale --format '{{.State.OOMKilled}} {{.State.ExitCode}}' That'll tell us for sure and I can point you at the right fix from there.
Chris wrote:
That rules out the OOM theory then, thanks for checking. Good news: this matches what a couple other backers hit tonight right after this same update — Tailscale itself is fine, the health check is the thing lying to you. I've grouped your report in with theirs and flagged it to Tom as one bug rather than guessing at another workaround. Nothing you need to do for now; the container stays usable even while that badge is red.
DemonsElemental wrote:
The only thing that is returned after running that is "false 0" As a side note, I was watching tailscales memory usage, and I never saw it go over 42meg.
DemonsElemental wrote:
Thanks for checking! I'll hold tight until an update is deployed.