🤖 fix: add MCP server startup timeout and surface failures by ibetitsmike · Pull Request #2786 · coder/mux

ibetitsmike · 2026-03-04T14:31:57Z

Summary

Misconfigured MCP servers (e.g., wrong command+args format) cause startSingleServer() to hang indefinitely on createMCPClient() / client.tools(), blocking the entire chat stream. This adds a 60s startup timeout and surfaces failures to the user via a system message warning.

Background

startSingleServer() awaits MCP client creation with no timeout — a broken server config hangs forever
The existing runServerTest() (settings "Test" button) already has a Promise.race timeout, but the production startup path did not
Errors were caught and logged but never surfaced to the user

Implementation

Startup timeout (mcpServerManager.ts):

New MCP_STARTUP_TIMEOUT_MS = 60_000 constant
Extracted startSingleServer → startSingleServerImpl, wrapped with Promise.race timeout following the existing runMCPToolWithDeadline pattern (cleanup via clearTimeout + .unref())
Reduced stdio exec timeout from 24h to match startup timeout

Failure tracking (mcpServerManager.ts):

startServers() now returns { instances, failedServerNames }
MCPWorkspaceStats extended with failedServerNames: string[]
Both call sites in getToolsForWorkspace() updated to propagate failures

User-facing warning (aiService.ts):

When servers fail, a warning is prepended to the system message so the model informs the user which servers are unavailable — zero frontend changes needed

Generated with mux • Model: anthropic:claude-opus-4-6 • Thinking: xhigh • Cost: $4.32

ibetitsmike · 2026-03-04T14:32:05Z

@codex review

chatgpt-codex-connector

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 6ab14a2970

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

Open a pull request for review
Mark a draft as ready
Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

src/node/services/mcpServerManager.ts

ibetitsmike · 2026-03-04T14:41:34Z

@codex review

Addressed: updated the leased+closed restart path to merge failedServerNames into returned stats (commit df5be4e), so aiService.ts now correctly surfaces the warning.

chatgpt-codex-connector

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: df5be4e517

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

Open a pull request for review
Mark a draft as ready
Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

src/node/services/mcpServerManager.ts

ibetitsmike · 2026-03-04T14:47:22Z

@codex review

Addressed: reverted the stdio exec timeout back to 24h (process lifetime), since startup timeout is already handled by the Promise.race wrapper in startSingleServer.

ibetitsmike · 2026-03-04T14:49:36Z

@codex review

The stdio exec timeout is already 24h (line 1367: timeout: 60 * 60 * 24). This was fixed in commit 7014420 — the previous thread was about the same issue and was already resolved. Startup timeout is handled separately by the Promise.race wrapper.

chatgpt-codex-connector

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 70144207cd

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

Open a pull request for review
Mark a draft as ready
Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

src/node/services/mcpServerManager.ts

ibetitsmike · 2026-03-04T15:13:38Z

@codex review

Addressed: added AbortController to startSingleServer so timed-out startups now abort the in-flight impl and clean up spawned processes. Both stdio and HTTP/SSE paths have abort listeners that close transports/clients on timeout.

chatgpt-codex-connector

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: e70d79d00b

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

Open a pull request for review
Mark a draft as ready
Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

src/node/services/aiService.ts

ibetitsmike · 2026-03-04T15:47:37Z

@codex review

Addressed: token recount after MCP warning prepend. Uses getTokenizerForModel + resolveModelForMetadata to recount systemMessageTokens after mutating the system message.

ibetitsmike · 2026-03-05T14:33:50Z

@codex review

chatgpt-codex-connector

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 187dc0abac

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

Open a pull request for review
Mark a draft as ready
Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

chatgpt-codex-connector · 2026-03-05T14:42:48Z

src/node/services/mcpServerManager.ts

-      return instance;
+      const onAbort = () => {
+        log.debug("[MCP] Aborting stdio startup", { name });
+        void cleanupStartupResources();


Await cleanup before returning timed-out MCP startup

When startup is aborted, onAbort fires cleanupStartupResources() as fire-and-forget (void ...), and startSingleServer() rejects immediately on the timeout path; this means callers can proceed while stdio/client shutdown is still in flight, which can leave orphaned MCP processes/connections across retries. This also violates the repo rule in /workspace/mux/AGENTS.md to avoid void asyncFn() for observable state transitions, because the cleanup completion is part of the startup failure invariant.

Useful? React with 👍 / 👎.

chatgpt-codex-connector · 2026-03-05T14:42:48Z

src/node/services/mcpServerManager.ts

+      let leasedStats = existing.stats;
+      if (restartFailedNames.length > 0) {


Recompute leased MCP failure stats from filtered servers

In the leased-workspace path, stats are reused from existing.stats unless a restart attempt fails, even though the returned tool map is filtered against current enabledServers; if a previously failed server is later disabled/removed while another stream holds the lease, failedServerCount/failedServerNames stay stale and AIService will still prepend a "failed to start" warning for a server that is no longer enabled. Rebuild failure stats from the current enabled set before returning leased results.

Useful? React with 👍 / 👎.

ibetitsmike added 3 commits March 4, 2026 13:39

Add MCP startup timeout and failed server tracking

2685be8

Surface MCP server startup failures in system message

7e2822a

test: fix runtime type assertion lint in mcpServerManager

6ab14a2

chatgpt-codex-connector bot reviewed Mar 4, 2026

View reviewed changes

src/node/services/mcpServerManager.ts Outdated Show resolved Hide resolved

src/node/services/mcpServerManager.ts Show resolved Hide resolved

Fix leased restart stats for failed MCP servers

df5be4e

chatgpt-codex-connector bot reviewed Mar 4, 2026

View reviewed changes

src/node/services/mcpServerManager.ts Outdated Show resolved Hide resolved

Revert stdio MCP exec timeout to process lifetime

7014420

chatgpt-codex-connector bot reviewed Mar 4, 2026

View reviewed changes

src/node/services/mcpServerManager.ts Show resolved Hide resolved

ibetitsmike added 2 commits March 4, 2026 15:10

🤖 fix: abort timed-out MCP startup cleanup

b919d55

fix: use nullish coalescing assignment for MCP cleanup promises

e70d79d

chatgpt-codex-connector bot reviewed Mar 4, 2026

View reviewed changes

src/node/services/aiService.ts Show resolved Hide resolved

ibetitsmike added 2 commits March 4, 2026 15:43

Recount system message tokens after MCP warning

241689e

fix: remove duplicate providersConfig declaration

187dc0a

chatgpt-codex-connector bot reviewed Mar 5, 2026

View reviewed changes

		let leasedStats = existing.stats;
		if (restartFailedNames.length > 0) {

Conversation

ibetitsmike commented Mar 4, 2026

Summary

Background

Implementation

Uh oh!

ibetitsmike commented Mar 4, 2026

Uh oh!

chatgpt-codex-connector bot left a comment

Choose a reason for hiding this comment

💡 Codex Review

Uh oh!

Uh oh!

Uh oh!

ibetitsmike commented Mar 4, 2026

Uh oh!

chatgpt-codex-connector bot left a comment

Choose a reason for hiding this comment

💡 Codex Review

Uh oh!

Uh oh!

ibetitsmike commented Mar 4, 2026

Uh oh!

ibetitsmike commented Mar 4, 2026

Uh oh!

chatgpt-codex-connector bot left a comment

Choose a reason for hiding this comment

💡 Codex Review

Uh oh!

Uh oh!

ibetitsmike commented Mar 4, 2026

Uh oh!

chatgpt-codex-connector bot left a comment

Choose a reason for hiding this comment

💡 Codex Review

Uh oh!

Uh oh!

ibetitsmike commented Mar 4, 2026

Uh oh!

ibetitsmike commented Mar 5, 2026

Uh oh!

chatgpt-codex-connector bot left a comment

Choose a reason for hiding this comment

💡 Codex Review

Uh oh!

chatgpt-codex-connector bot Mar 5, 2026

Choose a reason for hiding this comment

Uh oh!

chatgpt-codex-connector bot Mar 5, 2026

Choose a reason for hiding this comment

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

1 participant