Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Known limitations and untested behaviour

From cognokratos/simple-agent-template · docs/LIMITATIONS.md · pinned revision c66ce19d7b0c

Stated rather than implied. A control that is documented but unverified is worse than one that is absent, because it is believed.

Not tested automatically

BehaviourWhy notHow to check by hand
Nonce conflict under real concurrencyNeeds a live PostgreSQL; the constraint is a primary key, enforced by the databaseTwo concurrent spends of one approval token against a running cluster
Rollback of a failed audit insertSameBreak the audit insert and confirm the status is unchanged and the nonce free
End-to-end approval through the browserNeeds a cluster, a model that calls the function, and a humanEnable the feature and follow APPROVALS.md
Keycloak login through a real browserNeeds the clustermake dev, then sign in
The evaluation suites' actual scoresNon-deterministic and model-dependentmake eval-all with a model available
Trace export reaching MLflowNeeds the cluster and a modelmake trace-test

The make targets above exist and are documented; they are simply not part of any automated gate.

Deliberate gaps

A fabricated prior user turn is not screened. The input rail screens the latest turn and any client-supplied assistant turn. It does not re-screen prior user turns, because doing so made one refusal poison the rest of a conversation. The same caller can send that text as the latest turn, where the full rail does screen it. See GUARDRAILS.md.

PII masking costs streaming. NeMo's streaming rail runner can only use an action's result to decide blocked/not-blocked, never to rewrite text, so while mask sensitive data on output is enabled the middleware buffers the complete answer, masks it once and only then releases it (TextGuardrailsMiddleware._stream_with_buffered_masking; see GUARDRAILS.md). Answers are masked, but no longer stream token by token, and one over GUARDRAILS_PII_MAX_BUFFER_CHARS is refused. The configured score_threshold is also not honoured by the pinned release's masking action — the effective floor is Guardrails' hardcoded 0.4. Both are asserted by verify_output_guardrails.py so they cannot drift unnoticed.

Header redaction is not content redaction. The telemetry processor removes credential-bearing headers. A secret inside a tool result or a model answer is not reached by it. See OBSERVABILITY.md.

NAT's own identity_header refusal is advisory on the workflow routes. Configured, NAT 1.9 raises IdentityHeaderError for a missing, empty or repeated identity header and registers a handler that would answer 401. That handler is not reached: add_generate_routes serves the workflow path and its /stream and /full variants through the interactive runner unconditionally, and that runner acquires the session in a background task wrapped in a blanket except Exception, so the caller gets 200 with a WORKFLOW_ERROR in the stream. RequireIdentityHeaderMiddleware in fastapi_worker.py is what actually enforces the requirement, and make auth-test asserts it. See SECURITY.md.

Per-user trace attribution is off by default. NAT 1.9 stamps every span with the authenticated user (user.id and nat.user.id). That is genuinely useful for triage, and it is withheld unless OTEL_TRACE_USER_ID=true, because the traces already carry the question and the answer — the identifier is what turns them from a corpus into a per-person record, and whether that is acceptable depends on the trace store's access controls and retention. The value is a stable uuid5 pseudonym rather than the Keycloak subject, which is a weaker disclosure but not anonymity: it is the same value for the same person on every request. UserIdentityProcessor in observability/trace_processor.py. The raw gateway identity headers NAT copies into span metadata are redacted in both modes, so the switch governs the only per-user identifier a trace can carry.

Sessions are in memory. One gateway instance, and a restart logs everyone out.

An interaction with no recorded owner is allowed through unless HITL_STRICT_INTERACTION_OWNERSHIP=true, so NAT's own OAuth consent flow keeps working. Every interaction the approval module creates is recorded.

Resource requirements

make verify-output-guardrails loads Presidio's analyzer, which pulls spaCy's en_core_web_lg into memory — roughly 600 MB on top of the agent's own footprint. On a Docker VM already near capacity the kernel kills it, which surfaces as a bare exit 137 rather than a failing assertion. The script warns before that point. Give Docker headroom, or run the same script on the host where the dependencies are installed.

Observed: on a 7.7 GB Docker VM with MLflow at 2 GB and an unrelated stack running, the masking half was OOM-killed while the configuration, pattern and wiring halves passed. The same script passed in full on the host.

This is not confined to the verification script. Any live request whose answer reaches the mask sensitive data on output flow loads the same analyzer, so on a VM without that headroom the agent process is SIGKILLed mid-stream while masking. It leaves no Python-level error — the client sees the intermediate-step events, then a truncated stream (curl: (18)), and the container restarts with RestartCount incremented, OOMKilled=false and exit code 0, none of which name memory as the cause. make trace-test fails as "no streamed data chunks were returned".

Measured on the same 7.7 GB VM: with MLflow running the request was killed every time; stopping MLflow alone (freeing ~2 GB) made the same request return its masked answer with no restart. If make trace-test fails that way, check docker inspect <agent> --format '{{.RestartCount}}' across the request before looking for a fault in the agent.

Dependency constraints

nvidia-nat-security[guardrails]==1.9.0 pins nemoguardrails>=0.11,<0.22, so 0.23.0 — which fixes three streaming rail defects — cannot be installed. guardrails_compat.py works around them from application code and self-disables once the installed release is correct. Delete it when the pin allows >=0.23.

The 1.9 upgrade did not relax this. The requirement is byte-identical to 1.8.0's. So are nemo_guardrails_middleware.py, execution_store.py, routes/execution.py and nat/llm/openai_llm.py, and the ReAct _stream_fn still buffers until it sees Final Answer:. Every workaround in agent/src/nat_streaming_react/ therefore still has a reason to exist after the upgrade; none became deletable. See EXTENDING.md for the per-module removal conditions.

The observability package relies on three private NAT attributes, each listed with its removal condition in observability/__init__.py and OBSERVABILITY.md. This is not a purely public-API implementation, and all three are still private in 1.9.

Before production

This is a local demonstration. Add:

  • authorization and tenant/user scoping in every SQL query — the MCP tools currently return any row the query matches;
  • secrets management instead of the demo credentials in docker-compose.yml;
  • database migrations rather than a one-time init script;
  • pagination and response-size limits for history-heavy records;
  • access controls, retention and redaction for OpenTelemetry and MLflow data;
  • a dedicated low-latency guard model rather than sharing the application LLM;
  • explicit image digest pinning and vulnerability scanning;
  • a session store that survives a restart and supports more than one instance.

Licensing

Original code and documentation are licensed under MIT (root LICENSE). The package metadata (gateway/Cargo.toml, mcp-server/Cargo.toml, ui/package.json) and the SPDX headers of the original Python sources say the same.

Three files under agent/src/nat_streaming_react/ are exceptions. register.py and text_guardrails.py are modified from NVIDIA NeMo Agent Toolkit code, and observability/otlp_exporter.py closely follows it. They keep their Apache-2.0 declarations and NVIDIA's copyright notices, which is why agent/pyproject.toml declares MIT AND Apache-2.0. The list and the Apache-2.0 text are in THIRD_PARTY_NOTICES.md and LICENSES/Apache-2.0.txt. When you fork, keep those notices with the files they cover.

The nat-streaming-react distribution built from agent/ (and the agent image) carries the licence documents too. agent/LICENSE, agent/LICENSES/Apache-2.0.txt and agent/THIRD_PARTY_NOTICES.md are byte-identical copies of the root files, which stay authoritative, and are listed in project.license-files. make license-check fails if a copy drifts. make package-license-check builds the sdist, the wheel, a wheel from the sdist and an installed copy, and checks each for the complete texts.