Adding Tools & Integrations — Definition of Done
Use this checklist whenever you add or materially change:- a tool — under
integrations/<vendor>/tools/for a single-vendor tool, ortools/system//tools/cross_vendor/for a cross-cutting one (see tool-placement-policy.md) - an integration under
integrations/<name>/— its config, client, verifier, and tools - investigation source wiring for an existing tool or integration
1. Tool checklist
Files usually involved
integrations/<vendor>/tools/<tool_name>_tool/__init__.py— the tool package (most common path: the tool belongs to a vendor integration)tools/system/<tool_name>/ortools/cross_vendor/<tool_name>/— only when the tool is not vendor-specific (e.g.tools/system/sre_guidance_tool/)integrations/<name>/client.py— reuse a dedicated integration API client instead of inlining requestscore/tool_framework/utils/— shared helper code reused across vendorsdocs/<tool_name>.mdx— user-facing usage, parameters, examplestests/tools/test_<tool_name>.py— behavior and regression coverage
tools/ and integrations/<vendor>/tools/, so placement is about ownership, not discovery — see tool-placement-policy.md. Wherever a tool lives, it calls integration-local clients/helpers rather than inlining transport, and never lives in a top-level vendors/ or services/ package.
Tool packages must be substantive production modules — no empty or discovery-only __init__.py, no thin wrapper that only satisfies registry import. Any tool with validation, credential/parameter resolution, transport/client calls, output normalization, or error handling should split those concerns into focused sibling files (tool.py, models.py, validation.py, delivery.py/client.py, results.py), leaving __init__.py as a small registry entrypoint that imports the public tool object.
Contract and implementation
- Pick the simplest shape that fits (
@tool(...)for lightweight tools, a richer class only when needed) -
__init__.pyis a small registry entrypoint; non-trivial tools use sibling modules for implementation concerns - Metadata is complete and accurate:
name,description,source,surfaces,requires, and anyuse_cases/outputs/retrieval_controls -
input_schemamatches the actual runtime arguments and required fields -
is_availablereturnsTrueonly when the tool can genuinely run -
extract_paramsmaps resolved integration state into tool args correctly - Validation, credential/parameter resolution, transport/client calls, and result formatting are separated so each can be tested independently
- Reusable transport or integration-specific parsing lives in
integrations/<name>/orcore/tool_framework/utils/, not copied into the tool body - Failure responses have a stable, investigation-friendly shape; expected external failures (missing config, auth, rate limit, upstream 4xx/5xx) return structured errors rather than raising — unexpected exceptions use the global
BaseToolwrapper intentionally or are migrated with telemetry coverage - Output is normalized enough for the planner/LLM to consume reliably
- Secrets never leak through
extract_params, return values, logs, or traceable tool-call kwargs; secret/PII output is run throughinfrastructure/safety/masking/before return - External side effects declare
side_effect_level,requires_approval, andapproval_reasonwhere appropriate - To appear in both investigation and chat, set
surfaces=(ToolSurface.INVESTIGATION, ToolSurface.CHAT)
Live payload parsing
If the tool parses API, MCP, log, or webhook payloads:- Validate against the real or documented upstream response shape, not only idealized mocks
- Handle alternate field names used in live payloads
- Handle missing or partial fields without returning unusable output
- Preserve important context when truncating, tailing, paginating, or flattening data
- Upstream 429 / 5xx responses return a clear, investigation-friendly error rather than raising
- Add at least one regression test using a realistic fixture payload
hasMore / cursor mismatches; content-vs-pointer shapes (logs_content vs logs_url-style payloads).
Skill guidance (optional)
A tool can carry workflow guidance the model reads on every call by shipping aSKILL.md. The guidance is appended to the tool’s description under a Workflow guidance: heading — it is permanent schema text, not a side channel. All guidance targeting one tool is combined and truncated at 2400 characters (tools/registry_skill_guidance.py), so budget it like description text: the longer the guidance, the more of every request it consumes.
Skill guidance and a harness playbook are independent and composable — a tool may have both. Skill guidance rewrites one tool’s description; it does not replace agent- or harness-level guidance for a multi-step flow. The GitHub tools (github_cli, ci_fix, security_fix) each ship a SKILL.md; adding one never means removing broader guidance.
When to add (two independent axes — evaluate both):
Neither is the default for most of ~70
tools/ packages. Missing SKILL.md is usually correct. Never add one-per-vendor stubs, or one skill per observability vendor when the failure mode is shared (query hygiene is one class, not Datadog + Grafana + CloudWatch copies).
Harness authoring template: core/agent_harness/prompts/skills/_template/SKILL_TEMPLATE.md.
File. A SKILL.md with YAML frontmatter and a markdown body:
- Explicit: add the file’s path to
_skill_guidance_files()intools/registry_skill_guidance.py. ASKILL.mdthat exists on disk but is absent from that tuple is never loaded — unlisted and missing files are skipped with no diagnostic, the same trap as forgetting adocs.jsonentry. - Or place it under
tools/system/python_execution_tool/skills/*/SKILL.md, which is discovered automatically.
unknown_tool— a name undertools:matches no registered tool; that target is dropped.invalid_metadata— missingname/description/tools, an over-longname/description, or anamethat is not lowercase kebab-case; the whole skill is skipped.parse_failed— malformed frontmatter YAML.
2. Integration checklist
Files usually involved
integrations/<name>/__init__.py— package facade: a docstring, plus re-exports of the public API when callers need them (see the__init__.pyrule in AGENTS.md)integrations/<name>/config.py— config model,classify(), validators, selectors, normalization helpersintegrations/<name>/client.py— a dedicated API client, when the integration makes direct remote callsintegrations/<name>/verifier.py— local verification logicintegrations/<name>/tools/<tool_name>_tool/— the vendor’s agent-callable tools (see §1)integrations/<name>/background_adapter.py— only for messaging integrations you want selectable as a background-RCA completion channel (see Notification channels)integrations/catalog.py— resolve the integration into the shared runtime configintegrations/verify.py— wire the local verification pathdocs/<name>.mdx— user-facing setup, usage, verificationtests/integrations/test_<name>.py, plustests/tools/,tests/e2e/, ortests/synthetic/where tools or scenarios exercise it
integrations/<name>/ owns everything about one vendor — config, resolution, clients, verifiers, helpers, and its tools. Only vendor-less (tools/system/) and cross-vendor (tools/cross_vendor/) tools live under top-level tools/.
Examples from the repo
- Datadog:
integrations/datadog/(withintegrations/datadog/tools/),integrations/catalog.py, tests undertests/integrations/datadog/andtests/tools/test_datadog_*.py. - Grafana:
integrations/grafana/(withintegrations/grafana/tools/),integrations/catalog.py,surfaces/cli/wizard/local_grafana_stack/, tests undertests/integrations/grafana/andtests/tools/test_grafana_*.py. - Hermes:
integrations/hermes/(withintegrations/hermes/tools/hermes_logs_tool/and.../hermes_session_evidence_tool/),surfaces/cli/commands/hermes.py,tests/hermes/,tests/synthetic/hermes/. - Bitbucket:
integrations/bitbucket/shows the facade layout — a one-line__init__.pybesideconfig.py,client.py,verifier.py, andtools/.
Core completeness
- Config, normalization, validators, and
classify()are in place underintegrations/<name>/config.py, leaving__init__.pya facade - Catalog resolution / env loading is wired correctly
- Verification path is wired in
integrations/verify.pyand adapters/registry as needed - Integration-local client added under
integrations/<name>/client.py(only if it makes direct remote calls) - Tool layer is wired and stable
- CLI setup flow is updated if the integration is user-configurable locally
- Background-RCA delivery is wired, or intentionally out of scope (see Notification channels)
-
opensre integrations setup <name>parity is added, or intentionally documented as out of scope - New required env vars / credentials are added to
.env.example(never.env) - Sensitive credentials follow the Credential resolution contract below
-
make verify-integrationspasses
Notification channels
Only if the integration should appear in/background notify set <channel> as a destination for a completed background RCA. See background-investigations.mdx for the user-facing behaviour.
Add integrations/<name>/background_adapter.py with three members:
- Register the adapter object in
bootstrap/adapters.py. Nothing auto-discovers it, and importing the module is not enough: imports are cached, so a re-import after the registry is cleared runs no module body. -
delivernever raises. Return"sent","failed: <reason>", or"missing <name> integration: <what to configure>". The string is persisted on the record and shown by/background show, so redact any credential in the reason. - Import the vendor client inside
deliver, not at module scope. These modules are imported when a user runs/background notify set, so a module-level client lands on that path. - Send the bounded summary from
infrastructure.delivery.notifications.rca_summary, not the full report, for any channel with a message-size limit.
/background notify set derives its allowed list from the registry, so there is no channel list to edit.
Credential resolution
Secret env names and non-secret config follow different write/read paths. Keep this contract when adding or changing an integration.
Hard rules for new code
- Never use bare
os.getenvfor a secret env name (*_TOKEN,*_KEY,*_PASSWORD,*_SECRET, connection strings, and similar). Useresolve_env_credentialfromconfig.llm_credentials(env first, then the credentials file). - Webhook /
*_URLvalues are never written to the credentials file (wizard routes them to store/.env, notsync_env_secret). Read with store → plainos.getenvonly. Webhook URLs often embed a secret token — treat them like passwords for logging/masking. - Leave
load_env_integration_servicesplain-env-only (startup-safe; no credentials-file read at boot). - Store still wins in
resolve_effective/ merge — env/credentials-file is the fallback tier only. - Tools receive credentials through
extract_params(resolved integration state), never their own env reads. At execution, keys listed in the tool’sinjected_paramsoverride model-supplied values, so the verified source wins even when the model passes a token. A tool’s resolver may read env only as the final fallback when nothing was injected —integrations/github/tools/github_cli/credentials.pyis the reference: explicit/injected token first, thenGITHUB_MCP_AUTH_TOKEN, thenGITHUB_TOKEN/GH_TOKEN. - Set
OPENSRE_DISABLE_KEYRING=1to skip local-file reads/writes (env and store still work).
resolve_env_credential (env → credentials file), sync_env_secret / save_credential (secret writes to the credentials file and .env), sync_env_values (.env keys, including secrets).
3. Investigation wiring
If the tool/integration is relevant to investigations:- Review alert-source seeding in
core/domain/alerts/alert_source.py - Review source-priority/prompt mapping in
tools/investigation/stages/gather_evidence/prompt.py - Review evidence/source registration in
core/domain/types/or related state models - Declare
@tool(evidence_mapper=...)(or set theevidence_mapperclass attribute on aBaseTool) if the tool’s output should be citeable in the report. The mapper lifts the raw output into the canonical report keys the evidence catalog cites, and it lives with the tool in its own package — the investigation stage stays vendor-agnostic. A tool with no mapper keeps its raw output only; the catalog never mints a citeable key for it. The coverage guard (tests/tools/investigation/stages/gather_evidence/test_evidence_mapper_coverage.py) fails until you either add the mapper or record the tool inevidence_mapper_baseline.txtas a deliberate known gap. - Add scenario coverage proving the tool surfaces useful RCA evidence
alert_source, review the source-to-tool maps explicitly.
4. Discovery and edge cases
For tools that list, search, or inspect resources:- Folder/nested resource layouts are considered where the upstream supports them
- Large result sets are capped or paginated intentionally
- Partial fetches are surfaced clearly (
truncated,fetch_error, etc.) - Time/order-sensitive results preserve causal ordering where it matters
5. Docs and tests
Docs
- Ship or update a
docs/page/section in the same PR (new tool, CLI command, pipeline behavior, or integration; and whenever a tool’s API/schema or an integration’s setup changes) - Any new
docs/page is registered indocs/docs.json(without the.mdxsuffix) so Mintlify navigation shows it - Investigation LLM tool-calling changes follow investigation-tool-calling.md
Tests
- Unit tests for config/normalization
- Tool contract tests, or equivalent schema/metadata coverage
- A registry/discovery test proves the tool is visible on the expected surface(s)
- Runtime behavior tests for success and failure paths
- At least one realistic fixture for live-payload parsing when external payloads are involved
- If investigation-relevant, a test proves the planner/agent can discover or invoke the tool through the normal runtime path (plus synthetic/scenario coverage when the loop depends on it)
-
tests/integrations/updated when integration wiring changes
Final gate (new integrations)
Everything above is complete, and:- Screenshot or demo GIF showing the integration working end-to-end
- E2E or synthetic test added
- CI checks pass (see CI.md)