envoyproxy/envoy

OutlierDetection decrements ejectTimeBackoff too quickly

オープン

#21,142 opened on 2022/05/04

 (3 件のコメント) (0 件のリアクション) (0 人の担当者)C++ (5,373 件のフォーク)batch import
area/outlier_detectionhelp wanted

Repository metrics

Stars
 (27,997 個のスター)
PR merge metrics
 (PR metrics pending)

説明

Title: OutlierDetection decrements ejectTimeBackoff too quickly

Description:

On every tick of the intervalTimer value of ejectTimeBackoff will be decremented for all not ejected hosts.

if (!host->healthFlagGet(Host::HealthFlag::FAILED_OUTLIER_CHECK)) {
  // Node seems to be healthy and was not ejected since the last check.
  if (monitor->ejectTimeBackoff() != 0) {
    monitor->ejectTimeBackoff()--;
  }
  return;
}

But during the interval there could be no evidence that host actually became healthy.

Repro steps:

Let's say we have a host that constantly returns 500 status codes and we configured detector for 3 consecutive 5xx and the interval is equal to 1s:

  1. Send 3 requests to the host, each request returns 500
  2. OutlierDetection ejects this host for baseEjectionTime * ejectTimeBackoff where ejectTimeBackoff is equal 1
  3. ejectTimeBackoff++
  4. OutlierDetection unejects host after baseEjectionTime
  5. After 1s ejectTimeBackoff is decremented
  6. Send 3 requests to the host again, each returns 500
  7. OutlierDetection ejects this host for baseEjectionTime * ejectTimeBackoff where ejectTimeBackoff is equal 1 again

Expected behaviour: step 7 should eject the host for baseEjectionTime * 2. So at the step 5 OutlierDetection decrements ejectTimeBackoff without any knowledge if host recovered or still unhealthy.

Conclusion

This behaviour is confusing and not really documented properly. I think ejectTimeBackoff should be decremented only when there is convincing evidence that host is recovered.

xref: https://github.com/envoyproxy/envoy/pull/14235

コントリビューターガイド