fix: update banner — conflict recovery path + server self-restart after update (#816)
* fix: update banner conflict recovery + server self-restart after update (#813 #814) * fix(update): restart must wait for in-flight update + reset force button on retry Two defects in the update banner flow found during review of PR #816: 1. Two-target race (webui + agent sequential) The client posts targets sequentially: webui succeeds and schedules a restart timer (2 s delay); client then posts agent; server begins agent fetch+pull; at T=2 s the restart timer fires os.execv mid-pull, killing the agent update and closing the client connection. User sees "Update failed (agent): Failed to fetch" even though webui did update, and the agent repo is in an unknown partial state. Fix: _schedule_restart() now blocks on _apply_lock before calling os.execv. If a second update is in flight when the timer fires, the restart thread waits until it completes. If nothing is in flight the lock acquire is instant, so no-op updates still restart immediately. 2. Stale force-update button across retries _showUpdateError sets btnForceUpdate to display:inline-block when res.conflict / res.diverged. Nothing resets it on the next retry, so a subsequent non-conflict error (e.g. network) leaves the stale force button visible pointing at the previous target. Fix: applyUpdates() now hides the force button and clears its data-target at the start of each attempt. Tests: - test_schedule_restart_waits_for_apply_lock: holds _apply_lock from a helper thread, verifies execv is delayed until the lock is released. - test_schedule_restart_still_fires_when_no_update_in_flight: sanity check that the common path still works with no contention. - test_apply_updates_resets_force_button_at_start: regression guard that the reset appears before the update loop begins. Full suite: 1683 passed, 0 failures. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(update): hold _apply_lock through execv + fix banner error layout Two fixes from Opus review: 1. TOCTOU gap in _schedule_restart (api/updates.py): the original pattern acquired _apply_lock, released it, then called os.execv — leaving a brief window where a new update could start between release and execv. Fixed by moving os.execv inside the 'with _apply_lock:' block so the process is replaced while still holding the lock; no new update can acquire it. 2. Banner CSS layout (static/index.html): #updateError was a direct flex child of .update-banner (display:flex row), so long error messages sat inline between #updateMsg and the buttons instead of below the message. Wrapped #updateMsg + #updateError in a flex-column container so errors stack vertically under the status line. * docs: add v0.50.134 CHANGELOG entry --------- Co-authored-by: nesquena-hermes <nesquena-hermes@users.noreply.github.com> Co-authored-by: Nathan Esquenazi <nesquena@gmail.com> Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -1469,6 +1469,14 @@ def handle_post(handler, parsed) -> bool:
|
||||
|
||||
return j(handler, apply_update(target))
|
||||
|
||||
if parsed.path == "/api/updates/force":
|
||||
target = body.get("target", "")
|
||||
if target not in ("webui", "agent"):
|
||||
return bad(handler, 'target must be "webui" or "agent"')
|
||||
from api.updates import apply_force_update
|
||||
|
||||
return j(handler, apply_force_update(target))
|
||||
|
||||
# ── CLI session import (POST) ──
|
||||
if parsed.path == "/api/session/import_cli":
|
||||
return _handle_session_import_cli(handler, body)
|
||||
|
||||
136
api/updates.py
136
api/updates.py
@@ -183,6 +183,111 @@ def check_for_updates(force=False):
|
||||
_check_in_progress = False
|
||||
|
||||
|
||||
def _schedule_restart(delay: float = 2.0) -> None:
|
||||
"""Re-exec this process after *delay* seconds.
|
||||
|
||||
Called after a successful update so that the freshly-pulled code is
|
||||
loaded on the next request, rather than running with a mix of old and
|
||||
new Python modules in sys.modules.
|
||||
|
||||
os.execv() replaces the current process image with a fresh interpreter
|
||||
running the same argv — sessions are preserved on disk, the HTTP port
|
||||
is reclaimed within the delay window, and the client's own
|
||||
``setTimeout(() => location.reload(), 2500)`` lands after the restart.
|
||||
|
||||
Coordinates with ``_apply_lock``: when the user updates both webui
|
||||
and agent, the client POSTs them sequentially. Without coordination
|
||||
the restart timer scheduled by the first update's success would fire
|
||||
while the second update's git-pull is still running, killing it mid-
|
||||
stream and leaving the second repo in an unknown partial state.
|
||||
Blocking on ``_apply_lock`` before ``os.execv`` means a pending
|
||||
second update always completes before the restart happens.
|
||||
"""
|
||||
import os
|
||||
import sys
|
||||
|
||||
def _do():
|
||||
import time
|
||||
time.sleep(delay)
|
||||
# Hold _apply_lock through os.execv so no new update can start between
|
||||
# the lock-release and the process replacement. Any in-flight update
|
||||
# finishes first (since it holds the lock), and then the process is
|
||||
# replaced while still holding the lock — meaning no new update can
|
||||
# sneak in during the brief TOCTOU window that existed with the
|
||||
# original acquire-release-execv sequence.
|
||||
# Threads die when execv replaces the process image, so the lock is
|
||||
# released atomically by the kernel.
|
||||
with _apply_lock:
|
||||
try:
|
||||
os.execv(sys.executable, [sys.executable] + sys.argv)
|
||||
except Exception:
|
||||
# Last-resort: if execv fails (e.g. frozen binary), just exit
|
||||
# so the process supervisor (start.sh / Docker) restarts us.
|
||||
os._exit(0)
|
||||
|
||||
threading.Thread(target=_do, daemon=True).start()
|
||||
|
||||
|
||||
def apply_force_update(target: str) -> dict:
|
||||
"""Force-reset the target repo to the latest remote HEAD.
|
||||
|
||||
Unlike apply_update() which requires a clean working tree and refuses
|
||||
merge conflicts, this discards all local modifications (checkout .) and
|
||||
resets to origin/<branch> — equivalent to what the diverged/conflict
|
||||
error messages ask the user to run manually.
|
||||
|
||||
Should only be called when apply_update() has already returned a
|
||||
response with ``conflict: True`` or ``diverged: True`` and the user
|
||||
has confirmed they want to discard local changes.
|
||||
"""
|
||||
if not _apply_lock.acquire(blocking=False):
|
||||
return {'ok': False, 'message': 'Update already in progress'}
|
||||
try:
|
||||
if target == 'webui':
|
||||
path = REPO_ROOT
|
||||
elif target == 'agent':
|
||||
path = _AGENT_DIR
|
||||
else:
|
||||
return {'ok': False, 'message': f'Unknown target: {target}'}
|
||||
|
||||
if path is None or not (path / '.git').exists():
|
||||
return {'ok': False, 'message': 'Not a git repository'}
|
||||
|
||||
_, fetch_ok = _run_git(['fetch', 'origin', '--quiet'], path, timeout=15)
|
||||
if not fetch_ok:
|
||||
return {
|
||||
'ok': False,
|
||||
'message': 'Could not reach the remote repository. Check your connection.',
|
||||
}
|
||||
|
||||
upstream, ok = _run_git(['rev-parse', '--abbrev-ref', '@{upstream}'], path)
|
||||
if ok and upstream:
|
||||
compare_ref = upstream
|
||||
else:
|
||||
branch = _detect_default_branch(path)
|
||||
compare_ref = f'origin/{branch}'
|
||||
|
||||
# Discard local modifications then reset to remote HEAD
|
||||
_run_git(['checkout', '.'], path)
|
||||
_, ok = _run_git(['reset', '--hard', compare_ref], path)
|
||||
if not ok:
|
||||
return {'ok': False, 'message': f'Force reset to {compare_ref} failed'}
|
||||
|
||||
with _cache_lock:
|
||||
_update_cache['checked_at'] = 0
|
||||
|
||||
_schedule_restart()
|
||||
|
||||
return {
|
||||
'ok': True,
|
||||
'message': f'{target} force-updated to {compare_ref}',
|
||||
'target': target,
|
||||
'restart_scheduled': True,
|
||||
}
|
||||
finally:
|
||||
_apply_lock.release()
|
||||
|
||||
|
||||
def apply_update(target):
|
||||
"""Stash, pull --ff-only, pop for the given target repo."""
|
||||
if not _apply_lock.acquire(blocking=False):
|
||||
@@ -235,7 +340,16 @@ def _apply_update_inner(target):
|
||||
# Fail early on unresolved merge conflicts
|
||||
if any(line[:2] in {'DD', 'AU', 'UD', 'UA', 'DU', 'AA', 'UU'}
|
||||
for line in status_out.splitlines()):
|
||||
return {'ok': False, 'message': 'Repository has unresolved merge conflicts'}
|
||||
return {
|
||||
'ok': False,
|
||||
'message': (
|
||||
f'The local {target} repo has unresolved merge conflicts. '
|
||||
'To reset to the latest remote version run: '
|
||||
'git -C ' + str(path) + ' checkout . && '
|
||||
'git -C ' + str(path) + ' pull --ff-only'
|
||||
),
|
||||
'conflict': True,
|
||||
}
|
||||
stashed = False
|
||||
if status_out:
|
||||
_, ok = _run_git(['stash'], path)
|
||||
@@ -296,4 +410,22 @@ def _apply_update_inner(target):
|
||||
with _cache_lock:
|
||||
_update_cache['checked_at'] = 0
|
||||
|
||||
return {'ok': True, 'message': f'{target} updated successfully', 'target': target}
|
||||
# Schedule a self-restart so the updated code is loaded fresh. A plain
|
||||
# git pull leaves stale Python modules in sys.modules — agent imports that
|
||||
# reference new symbols (functions, classes) added in the update will fail
|
||||
# on the next request with AttributeError / ImportError. os.execv() re-
|
||||
# execs the same interpreter with the same argv, picking up the new code
|
||||
# cleanly without requiring the user to restart manually.
|
||||
#
|
||||
# The 2 s delay gives the HTTP response time to flush to the client before
|
||||
# the process replaces itself. The client already does
|
||||
# setTimeout(() => location.reload(), 1500) on success, so the page reload
|
||||
# and the restart land at roughly the same time.
|
||||
_schedule_restart()
|
||||
|
||||
return {
|
||||
'ok': True,
|
||||
'message': f'{target} updated successfully',
|
||||
'target': target,
|
||||
'restart_scheduled': True,
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user