Fill protection and growth
DonkeyFleet measures destination logical usage against destination volume size. When usage reaches the automation profile's fill threshold, it opens one durable fill episode and creates a resize recommendation. The rule applies to DonkeyFleet-created relationships and adopted relationships that an operator explicitly bound to an automation profile. Unbound adopted and suspended relationships remain observation-only.
Two different resize paths
Do not confuse the autonomous step with the approved recommendation. They use different parameters and have different purposes.
| Path | Purpose | Target calculation | Authorization |
|---|---|---|---|
| Autonomous grow | Immediate breathing room | Current size increased by autonomous_grow_pct | Only auto mode, effective dry-run off, and every guardrail satisfied |
| Recommended resize | Bring the destination to the operating target | ceil(logical_used × 100 / target_utilisation) | Always requires human approval |
For example, a 100 GiB destination containing 86 GiB is 86% full. A 5% autonomous grow changes its size to 105 GiB and usage to about 81.9%. A recommendation targeting 70% changes its size to about 122.9 GiB. If the observed result is near 70%, the approved recommendation ran; the small autonomous step did not itself target 70%.
Decision flow
Parameters and their effects
| Automation parameter | What it controls | What it does not control |
|---|---|---|
reconcile_mode | Only auto permits an unprompted grow | It never removes the approval requirement for the full target resize |
fill_threshold | Opens an episode and produces the resize recommendation when logical usage is at or above this percentage | It does not determine either resize size |
autonomous_grow_pct | The one autonomous step — a bounded 5% or 10% choice added to the current destination size | It does not target target_utilisation |
target_utilisation | Calculates the human-approved recommended size | It does not size the autonomous step |
episode_reset_threshold | Closes an open episode only after usage falls below this percentage; this hysteresis prevents repeated proposals near the trigger | It does not close an episode merely because usage falls below fill_threshold |
max_cumulative_autonomous_growth_pct | Caps the sum of autonomous growth applied to that relationship across episodes | It does not cap a human-approved resize |
max_actions_per_run | Limits fresh autonomous grow submissions for each automation profile in one run; highest utilization is handled first | It does not stop recovery of an already persisted grow job |
The deployment-level effective dry-run is an additional gate. With dry-run on, DonkeyFleet can observe the breach and present the recommendation, but it submits no destination resize.
Snapshot-dominated destinations
Logical used includes active-file-system data and snapshots. If snapshots use more space than active data, DonkeyFleet holds the autonomous grow because retention may be the real cause. The notification shows both values so an operator can choose between resizing and changing retention. DonkeyFleet never changes snapshot retention automatically.
Capacity and safety
autonomous_grow_pct is a deliberately small, aware step: a bounded 5% or 10% choice (default
5%), enforced by validation. It only buys breathing room while an operator reviews the recommended
resize — the full resize to target_utilisation always requires human approval.
Before enabling auto mode:
- confirm aggregate headroom and any FSx SSD or capacity-pool limits;
- set a deliberate autonomous percentage and cumulative ceiling;
- test the episode in dry-run;
- verify the urgent, applied, and failed notifications reach operators.
Job and restart recovery
DonkeyFleet checkpoints the autonomous target before calling ONTAP and persists the returned job UUID. Later runs poll that job instead of issuing another write. If the process dies before the job UUID is stored, DonkeyFleet first checks whether the target size is already visible. A destination resize to a persisted absolute target is idempotent, so this one write may be safely resubmitted when its outcome remains indeterminate. PostgreSQL is therefore required recovery state.