# 備援做得太好，壞了一個月都沒人發現

- URL: https://justfly.idv.tw/%e5%82%99%e6%8f%b4%e5%81%9a%e5%be%97%e5%a4%aa%e5%a5%bd%ef%bc%8c%e5%a3%9e%e4%ba%86%e4%b8%80%e5%80%8b%e6%9c%88%e9%83%bd%e6%b2%92%e4%ba%ba%e7%99%bc%e7%8f%be/
- 日期: 2026-10-12
- 分類: 我知故我在
- 標籤: server, web, 備援, 可觀測性

![備援做得太好，壞了一個月都沒人發現]
備胎開起來跟原廠胎一樣順，就會忘了自己正用備胎跑高速公路。一個顯示即時匯率的功能，在線上就這樣跑了將近一個月。

##### 畫面一直有數字

數字在，格式對，沒有錯誤提示，沒有紅字。後來要加新功能，順手拿外部行情來比對，才發現線上回的根本是寫死的備援值。上游早已取不到資料，系統默默用預設值頂了將近一個月。

換算、分帳、顯示，全都建立在那組預設值上。每一層都運作正常，因為每一層收到的東西都長得像真的。

##### 分界點不在上游

外部來源本來就會掛，這不是問題。問題在備援值與真實值長得一模一樣，回應裡沒有任何欄位能區分「這是剛抓到的」還是「這是編的」。從使用者、前端到監控，沒有一層有機會察覺。

日誌裡其實留了一行「使用備援」。一行字躺在日誌裡，沒有人排程去讀它，也沒有任何告警接著它。等於備援啟動的事實只存在於一個沒人看的地方。

##### 無縫的備援，把可觀測性一起吃掉

直覺上，備援越無縫越好，使用者看不到錯誤就是好體驗。這個判斷在這裡正好反過來：[容錯](https://zh.wikipedia.org/wiki/%E5%AE%B9%E9%8C%AF)做得太乾淨，會把一個會被注意到的故障，換成一個不會被注意到的錯誤資料。故障至少會有人來報，錯誤資料只會被當成正常數字繼續用。

同一週還出現同型問題。某個定時快照排程因為主機重開，漏跑了幾輪，最後靠另一個「快照時間是否過舊」的告警才抓到。差別只在資料有沒有帶時間：有標時間的東西，過期了才有地方被發現。

##### 怎麼改，怎麼驗

備援改成有層次的鏈：主來源，次要來源，短時間內的舊[快取](https://zh.wikipedia.org/wiki/%E5%BF%AB%E5%8F%96)，最後才是寫死值。每一層都有逾時上限，資料不齊就整份退回下一層，半套資料不算成功。

每份回應再附上資料來源、抓取時間、報價日期、是否為備援、是否過期。上游報價日期太舊就不採用，舊報價也不得覆蓋新報價。客戶端同樣不盲信伺服器，依報價時間自行判斷是否標示「非即時」。

驗證只有一個動作：刻意切斷主來源，看畫面是不是真的顯示降級狀態。畫面若照常顯示數字，備援就還沒改好。

##### 留給之後的一句話

設計備援時，多問一句：它啟動了，誰會知道？答案若是「沒人」，那它不是備援，是消音器。降級狀態要寫進資料本身，而不是只寫在日誌裡。

下一次切斷主來源的驗證，排在新功能上線之前。

— 邱柏宇

### The Fallback Was So Good Nobody Noticed the Outage for a Month

A spare tire that drives exactly like the original is a spare tire easily forgotten while running on the highway. A feature showing live exchange rates ran that way in production for nearly a month.

##### The screen always had numbers

The numbers were there, formatted correctly, with no error message and nothing in red. Only while adding a new feature, and comparing against an external quote source on the side, did it turn out that production had been returning hardcoded fallback values. The upstream had been unreachable for a long time, and the system had quietly covered with defaults for almost a month.

Conversion, splitting, display: everything downstream was built on those defaults. Each layer behaved normally, because each layer received something that looked real.

##### The dividing line is not the upstream

External sources go down. That part is normal. The problem was that the fallback value looked identical to a real one. No field in the response could tell “freshly fetched” apart from “made up”. From the user to the frontend to monitoring, no layer had a chance to notice.

One line in the log did say a fallback was in use. A single line in a log nobody reads, with no alert attached, means the fact that the fallback kicked in lived only in a place nobody looked.

##### A seamless fallback swallows observability

The instinct says the smoother the fallback, the better, since users who see no error are having a good experience. Here the instinct is backwards. When [graceful degradation](https://en.wikipedia.org/wiki/Graceful_degradation) is too clean, it swaps a failure that would get noticed for wrong data that does not. A failure at least gets reported. Wrong data just gets used as if it were normal.

A same-shaped problem showed up the same week. A scheduled snapshot job missed several rounds after a host reboot, and it was caught only by a separate alert that checks whether the snapshot timestamp is too old. The difference came down to one thing: data carrying a timestamp can be found out when it goes stale.

##### The fix, and how to verify it

The fallback became a layered chain: primary source, then a secondary source, then a short-lived stale [cache](https://en.wikipedia.org/wiki/Cache_(computing)), and only last a hardcoded value. Every layer has a timeout, and incomplete data sends the whole payload down to the next layer. Partial data does not count as success.

Every response now carries its source, fetch time, quote date, whether it is a fallback, and whether it is stale. Upstream quotes that are too old are not used, and an old quote may not overwrite a newer one. The client does not trust the server blindly either, and decides on its own from the quote time whether to label the figure “not live”.

Verification is a single action: deliberately cut the primary source and check that the screen really shows a degraded state. If it still shows numbers as usual, the fallback is not finished.

##### One question for next time

When designing a fallback, ask once: if it kicks in, who will know? If the answer is nobody, it is a silencer, not a fallback. The degraded state belongs inside the data itself, not only in the log.

The next primary-source cutoff test is scheduled before the new feature ships.

— 邱柏宇
