Troubleshooting
Symptoms you are likely to meet on a first install, what each one actually means, and how to clear it.
Requests are denied#
| Response | Meaning | Fix |
|---|---|---|
| 403 policy_denied missing_policy |
Policy is enforcing and no allow rule matched. This is the default state of a fresh enforcing install. | Author policy.rules. See Write a policy that allows traffic. |
| 409 no_matching_agent | No healthy agent provides the requested capability. An agent in Degraded is deliberately not routed to. |
kubectl get agents -A, then inspect the pod behind the degraded agent. |
| 403 policy_denied tenant_denied |
The tenant is outside allowedTenants. |
Add the tenant, or send the request as one already allowed. |
| 409 approval_required | Working as intended: an approval rule matched and the request is waiting for a decision. | Approve or deny it in Console → Approvals. |
Pods will not start#
| Symptom | Meaning | Fix |
|---|---|---|
| CrashLoopBackOff on first install |
Services started before PostgreSQL was accepting connections. They retry the connect for up to a minute before exiting. | Usually clears itself. If it persists, the DSN or the database is genuinely wrong — read the pod logs. |
| connect: connection refused to the database |
The database is reachable by DNS but not accepting connections yet, or the port is wrong. | Confirm the database pod is ready and the DSN port matches the Service. |
| no such host | DNS could not resolve the database host, often a CoreDNS restart mid-install. | Check the hostname in the DSN and that CoreDNS is healthy. |
| ImagePullBackOff | The cluster cannot pull the images. | Confirm private-registry access and configure global.imagePullSecrets when your platform does not provide registry identity. |
| Migration Job fails | Checksums rejected the recorded history, meaning an applied migration was edited. | Do not edit applied migrations. Restore the database or roll forward with a new one. |
The gateway is not ready#
Read /readyz rather than guessing — it names the dependency that is failing:
kubectl port-forward -n agentfleet-system svc/agentfleet-gateway 8080:8080 &
curl -s http://127.0.0.1:8080/readyz
A required check reporting unavailable is what holds the gateway out of service. A check
reporting disabled is switched off in your values and is not a fault.
Helm upgrade fails#
Upgrading from an older release
If helm upgrade --reuse-values fails on a value that did not exist in the release you are
upgrading from, re-render with an explicit values file instead. --reuse-values does not merge
new chart defaults into an old release.