Ten minutes down on 29 September: what happened to WebAPI in Germany, and what changed overnight

One hung call to a trading server took the German WebAPI node down for ten minutes. Here is the timeline, the cause, and the fix that went to production the next morning: a deadline on every trading-server call, safe retries, and a failover that reacts in seconds.

Yesterday, 29 September, the WebAPI node in Germany (m4) stopped answering for about ten minutes. If your integration was routed there, you saw timeouts or connection errors between roughly 15:47 and 15:57 UTC. I'd rather you read what happened from me than guess from your logs.

What happened

All times UTC.

  • 15:28 — a trade request to one MT4 server went into the trading server and did not come back for 674 seconds. At 15:39 an MT5 request did the same for 348 seconds.
  • 15:47 — the node stopped answering everything, its own health checks included.
  • 15:48–15:50 — cloud.mywebapi.com did not answer for clients routed to Germany.
  • ~15:51 — the global router moved that traffic to our other regions. Requests worked again, just from further away.
  • 15:57 — the node was restarted automatically and came back.

If you call m4.mywebapi.com directly rather than cloud.mywebapi.com, you were cut off for the full ten minutes.

Why one slow call took everything down

Every call to a trading server waited for the answer with no time limit, and calls over one connection to a trading server run one at a time. So one request that hung inside the trading server made every other request to that server queue up behind it, and each waiting request held a worker. When the workers ran out, nothing else could be served — not even the health check that should have flagged the node as broken. And that health check was configured to tolerate several minutes of silence before anything acted on it.

That's the whole story: no single component crashed. A missing timeout turned one slow trading server into an outage for everyone on the node.

What changed — in production in every region since 30 September

Every trading-server call has a deadline. Trades 5 s, reads 10 s, changes 15 s, history and reports 30 s, server maintenance 60 s. If the trading server doesn't answer in time, you get an answer anyway, and you never wait minutes. You can set your own deadline per request with the X-Request-Timeout header (1–300 s) — shorter for a latency-sensitive screen, longer for a large report.

The answer tells you what a timeout means. A read that timed out changed nothing and is safe to repeat (Timeout). A trade or change that timed out may still be completed by the trading server (OutcomeUnknown) — so don't repeat it blindly. If too many requests are already waiting for one trading server, new ones are refused before they're sent (Busy) and are safe to repeat. v1 answers 504 / 503 with the same meaning in the X-Request-Outcome header.

A hung call no longer blocks the rest. A connection stuck inside the trading server is taken out of service and a fresh one is opened, so the next request goes through while the stuck one finishes on its own.

Repeating a trade is safe with an idempotency key. Send Idempotency-Key with a trade or change you might need to repeat. While the first request is still running, a repeat with the same key is not executed again; once it has finished, the repeat returns the original result — even if the first caller had already given up with OutcomeUnknown. This also closes an older gap: two requests sent at the same moment with the same key used to both execute.

Failover reacts in well under a minute. The global router now checks each region every 10 seconds and moves traffic away after the third failed check; with DNS caching, clients follow within about a minute instead of several. A node that stops answering is marked unhealthy after three failed health checks — within about half a minute — and restarted.

The details — defaults, headers, how to repeat a trade after OutcomeUnknown — are in the WebAPI documentation.

What you might want to do

  • Use cloud.mywebapi.com, not a regional host. It's what moves you to a healthy region when one isn't.
  • Keep your HTTP client's own timeout about 30 seconds longer than the deadline you send, so you receive our answer instead of cutting the connection yourself.
  • Send an Idempotency-Key with trades. It's the one thing that makes "did my trade go through?" answerable after a timeout.
  • Update the SDK to 0.3.0 — TypeScript (@mywebapi.com/sdk), .NET (MyWebApi.Sdk), PowerShell (MyWebApi) and, new, Python (mywebapi-sdk on PyPI). It sets a deadline per call, reports the new outcomes, and never repeats a trade or a change by itself. Older .NET versions don't know the new error codes, so update before you rely on them.

I'm sorry for the ten minutes. The fix isn't "we'll watch more closely" — it's that a hung trading server can no longer take the node down with it.