recourse¶
recourse (n.): a source of help or strength.
recourse coordinates resilience for Go services. Retries, hedging, circuit breaking, timeouts, and budgets work together under a shared policy, with structured visibility into execution.
Call sites provide a stable policy key. Local or remote policies define how calls handle failures, slow responses, and additional load. Teams can tune that behavior consistently across services without rebuilding resilience controls at every call site.
New here? Start with Design overview, then move to Getting started.
Why recourse?¶
A slow or failing dependency forces several decisions: whether to wait, start another attempt, reject the call, or limit additional work. Those decisions affect both latency and load.
recourse brings these controls into one execution path:
| What you need to control | What recourse provides |
|---|---|
| Transient failures | Retries with bounded attempts, backoff, and classifiers that decide which outcomes warrant another attempt. |
| Slow responses | Hedging that can start additional attempts after a fixed delay or a latency-based trigger. |
| Repeated dependency failures | Circuit breaking that rejects calls while the circuit is open and probes for recovery after a cooldown. |
| Additional load | Budgets that allow or deny retry and hedge attempts. |
| Time spent on a call | Per-attempt and overall timeouts configured through policy. |
| Incident diagnosis | Timelines and observer hooks that expose attempt outcomes, timing, and execution decisions. |
The controls also account for one another. When a circuit breaker probes for recovery, hedging is disabled so the probe does not launch extra requests. Configured budgets limit the additional work from retries and hedges.
What “policy-driven” means¶
Call sites supply a stable, low-cardinality key, such as "payments.Charge". The provider resolves a policy for that key, and the executor applies its configured behavior:
- Attempt limits, backoff, and jitter.
- Hedge limits and triggers for starting concurrent attempts.
- Circuit breaker thresholds and recovery cooldowns.
- Per-attempt and overall timeouts.
- Classifier selection to interpret errors and results.
- Budgets to gate retries and hedges.
Policies can live in your application through StaticProvider or come from an external source through RemoteProvider. The remote provider caches policies, coalesces concurrent fetches, and offers opt-in last-known-good fallback during source failures. See Remote Configuration.
This separates the operation from its resilience configuration: the call site keeps doing its work while the policy defines how that work is protected.
Quick start¶
The facade API takes a string key like "user-service.GetUser":
package main
import (
"context"
"github.com/aponysus/recourse/recourse"
)
type User struct{ ID string }
func main() {
user, err := recourse.DoValue[User](context.Background(), "user-service.GetUser", func(ctx context.Context) (User, error) {
// call dependency here
return User{ID: "123"}, nil
})
_ = user
_ = err
}
This minimal example uses the default policy, which provides bounded retries. Hedging, circuit breaking, timeouts, and budget limits are configured explicitly. See Getting started to install your policies and executor at startup.
When you need to know what happened, request a timeline:
ctx, capture := observe.RecordTimeline(ctx)
user, err := recourse.DoValue(ctx, "user-service.GetUser", op)
_ = user
_ = err
tl := capture.Timeline()
for _, a := range tl.Attempts {
// a.Attempt, a.Outcome, a.BudgetAllowed, a.Backoff, a.Err, ...
}
Observability-first¶
When a call fails or takes longer than expected, you need to know which policy ran, which attempts were launched, and what stopped further work.
recourse captures a structured observe.Timeline (attempt timings, outcomes, budget decisions, errors) and can also stream attempt/timeline events to your own logging/metrics/tracing via observe.Observer.
What’s inside¶
- Policy keys: stable, low-cardinality keys (
"svc.Method") that select behavior. - Policies + providers:
policy.EffectivePolicyresolved through in-processcontrolplane.StaticProvideror cachedcontrolplane.RemoteProviderlookups. - Execution:
retry.Executorcoordinates retries, hedging, circuit breaking, and timeouts according to policy. - Classifiers: pluggable
(value, err) → Outcomedecisions to retry, stop, or abort based on protocol and domain semantics. - Budgets/backpressure: gates on retry and hedge attempts to control additional load.
- Observability: structured
observe.Timelineplus streaming hooks viaobserve.Observer.
Where to go next¶
- Design overview – decision-first intro and tradeoffs.
- Getting started – install and first examples.
- Gotchas & safety checklist – avoid common operational failures.
- Adoption guide – staged rollout plan.
- Incident debugging – timeline-based runbook.
- Migrating from cenkalti/backoff – translate a familiar retry loop into governed retry policy.
- API compatibility policy – v1 stability contract.
- Defaults and safety model – generated defaults and failure modes.
- Policy schema reference – generated field reference.
- Reason codes & timeline fields – generated reference.
- Changelog – release history.
- Concepts:
- Policy keys
- Key patterns and taxonomy
- Policies & providers
- Classifiers
- Observability
- Budgets & backpressure
- Hedging
- Circuit Breaking
- Remote Configuration
- Integrations
- Architecture decisions:
- ADR 001: Low-cardinality policy keys
- ADR 003: Policy normalization
- Extending – write custom classifiers/budgets/observers.