Skip to main content

Chompy MCP HTTP API Documentation

Introduction​

The Chompy Model Context Protocol (MCP) HTTP API exposes your network observability platform — devices, flow, SNMP metrics, syslog, configuration history, correlated incidents, and IPAM — to AI agents, internal LLMs, and automation workflows over a simple authenticated REST surface.

This API is designed for agentic use. Beyond raw telemetry, it surfaces the platform's reasoning layer: correlated incidents with root-cause signals, ordered signal timelines, and pre-aggregated analytics (top conversations, application breakdowns, time-series). Many tools also return caveat fields — inline epistemic flags (stale, low_sample, high_retransmit, indefinite_windows_present, ordering_warning, …) that tell an agent how much to trust a result, so it reasons soundly instead of taking every number at face value.

Note on surfaces. This document describes the HTTP API (/api/mcp/tools/...), authenticated by an MCP API key. It is distinct from the stdio MCP server used by desktop LLM clients. Active-probe tools (run_ping_test, run_traceroute, run_http_test) currently live in the stdio server and are not part of this HTTP surface; they are listed in the roadmap below.


Table of Contents​

  1. Getting Started
  2. Authentication & Permissions
  3. API Reference
  4. Caveat Fields
  5. Tool Reference
  6. Code Examples
  7. Error Handling
  8. Rate Limits
  9. Best Practices
  10. Roadmap: Agentic Action Tiers

Getting Started​

Base URL​

https://your-chompy-instance.com/api/mcp

Quick Start​

  1. Obtain an MCP API key from your Chompy administrator (keys are prefixed chmp_).
  2. Make your first request:
curl -X POST https://your-chompy-instance.com/api/mcp/tools/network_health_summary \
-H "Authorization: Bearer chmp_your_api_key" \
-H "Content-Type: application/json"

All tool calls are POST requests with a JSON body of parameters (an empty {} is fine for tools that take none).


Authentication & Permissions​

All requests authenticate via a Bearer token (the MCP API key):

Authorization: Bearer chmp_your_api_key

Each key carries a permission level. Levels are hierarchical — higher levels include everything below them. A key may also be restricted to an explicit allowed_tools list, in which case it can call only those tools (within its level).

LevelIncludesUse for
read_onlyAll read tools (telemetry, incidents, analytics, config, logs, context, IPAM)Dashboards, monitoring agents, investigation/diagnosis
read_writeread_only + acknowledge_alertAgents that triage and acknowledge
adminread_write + query_postgres, query_clickhouseOperators needing raw SELECT access

Scoping agents. For an autonomous agent, issue a read_only key (optionally narrowed via allowed_tools) so it can observe and diagnose the entire estate but cannot mutate anything. Action capabilities are introduced separately — see the roadmap.

Management endpoints (/keys, /audit) require an admin session token, not an MCP API key — they are operator surfaces, not agent surfaces.


API Reference​

List available tools​

GET /api/mcp/tools

Returns every tool, the permission it requires, and whether your key can call it. Requires a valid MCP API key.

{
"tools": [
{ "name": "get_devices", "permission_required": "read_only", "available": true },
{ "name": "query_postgres", "permission_required": "admin", "available": false }
],
"your_permission": "read_only"
}

Execute a tool​

POST /api/mcp/tools/{tool_name}

Headers

Authorization: Bearer chmp_your_api_key
Content-Type: application/json

Body: tool-specific parameters (see Tool Reference).

Success response

{ "success": true, "data": { /* tool-specific */ } }

Error response

{ "error": "Description of what went wrong" }

Caveat Fields​

Many tools include a caveats array (and some include named flags like ordering_warning) in their data. These are machine-readable epistemic signals: they describe limitations of the result so an agent can weigh it correctly. Treat a populated caveat as a reason to qualify conclusions, gather corroborating evidence, or widen a time window.

CaveatMeaning
staleThe underlying data's most recent sample is older than expected (e.g. a device that has stopped being SNMP-polled). A down/flat reading may reflect missing data, not reality.
no_dataNo matching records in the window.
low_sampleToo few data points to establish a trend.
high_retransmitTCP retransmission rate is elevated — degraded transport.
high_rttAverage round-trip time is elevated.
high_port_fanoutA host contacted an unusually large number of distinct ports — possible scan/exfil.
cleanExplicitly nothing wrong found (e.g. no interface errors).
quietNo activity reached a threshold (e.g. no log bursts).
truncatedOutput was capped; more data exists.
indefinite_windows_presentOne or more maintenance windows have no end time and will suppress incidents until manually cleared.
ordering_warningIn an incident timeline, a symptom signal predates the designated root cause — causal ordering may be inverted. Do not assume the root cause produced earlier signals.

Tool Reference​

Conventions. Time windows are specified in minutes and are bounded server-side. limit values are clamped to safe maxima. IPs accept addresses or CIDR. All string inputs are validated; invalid input returns 400.

Devices & Health​

get_devices​

List inventory devices with live status pulled from recent SNMP polls.

NameTypeDefaultDescription
device_typestring–Filter by type substring (router, switch, firewall)
sitestring–Filter by site name substring
rolestring–Filter by role substring
limitnumber50Max results (≤1000)

Device status is three-state, derived from snmp_metrics.device_status:

ValueLabelMeaning
2upHealthy — reachable by both ICMP and SNMP
1degradedReachable by ICMP or SNMP, but not both (partial)
0downNot reachable by ICMP or SNMP
(no recent poll)unknownNo SNMP data in the lookback window

degraded is a real warning state and often precedes a full outage — do not treat it as either up or down.

Example

curl -X POST https://chompy.example.com/api/mcp/tools/get_devices \
-H "Authorization: Bearer chmp_xxx" -H "Content-Type: application/json" \
-d '{"device_type": "switch", "limit": 10}'

Response

{
"success": true,
"data": {
"devices": [
{ "id": 12, "name": "cat9300-sw01", "ip": "192.168.100.2",
"type": "switch", "role": "access", "site": "HQ",
"status": "up", "cpu": 14.2, "memory": 38.5 }
],
"count": 1
}
}

get_device_details​

Detailed record for one device, including its interfaces.

NameTypeRequiredDescription
device_namestringYesDevice name or management IP (substring match on name)

Returns device (name, ip, type, role, site, three-state status, cpu, memory) and an interfaces array.

network_health_summary​

Estate-wide situational summary. No parameters.

Returns total_devices, a device_status breakdown (using the three-state labels), active_alerts by severity, flow_stats (last hour), and synthetic_tests success rate.

{
"success": true,
"data": {
"total_devices": 24,
"device_status": [
{ "status": "up", "count": 1 },
{ "status": "degraded", "count": 1 },
{ "status": "down", "count": 20 },
{ "status": "unknown", "count": 2 }
],
"active_alerts": [ { "severity": "critical", "count": 1 } ],
"flow_stats": { "flows_last_hour": 285000, "bytes_last_hour": "45.2 GB" },
"synthetic_tests": { "successful": 142, "total": 150, "success_rate": "94.7" }
}
}

Alerts​

get_alerts​

Threshold-rule alert instances.

NameTypeDefaultDescription
severitystring"all"Filter by severity, or all
statestring"active"active (triggered/active), a specific state, or all
limitnumber25Max results (≤500)

acknowledge_alert — read_write​

Acknowledge an active alert.

NameTypeRequiredDescription
alert_idnumberYesAlert instance ID
acknowledged_bystringNoWho/what acknowledged (defaults to API)
notesstringNoAcknowledgment note

Incidents & Reasoning​

These tools expose the correlated incident layer produced by the Problem Assembler — the platform's reasoning output, not raw telemetry. This is usually where an agent should start an investigation.

get_incidents​

List correlated incidents (WO-YYYY-NNNNN) with their root-cause signal.

NameTypeDefaultDescription
statestring"open"open, acknowledged, resolved, or all
severitystring"all"Filter by severity
minutesnumber–Restrict to incidents opened in the last N minutes
limitnumber25Max results (≤200)

Each incident includes incident_id, headline, state, severity, confidence, threat_type, signal_count, correlation_method, the root_cause_signal (type, family, detected_at), and affected device/site IDs.

get_incident_details​

Full record for one incident, including AI-synthesized root cause, blast radius, remediation steps, and fault_analysis (path/topology) when present.

NameTypeRequiredDescription
incident_idstring | numberYesThe WO-… ID or numeric problem ID

get_incident_timeline​

The ordered signals that compose an incident, each with its role (root_cause / symptom), timing, severity, and key metrics. Includes the ordering_warning caveat when a symptom predates the root cause.

NameTypeRequiredDescription
incident_idstring | numberYesThe WO-… ID or numeric problem ID
{
"success": true,
"data": {
"incident_id": "WO-2026-00469",
"ordering_warning": "Root-cause signal occurs after 1 symptom(s) in time — causal ordering may be inverted...",
"signals": [
{ "signal_id": 880, "role": "symptom", "type": "ping_loss_spike",
"detected_at": "2026-06-26T05:52:00Z", "device": "edge-1", "z_score": 6.1 },
{ "signal_id": 884, "role": "root_cause", "type": "top_talker_spike",
"detected_at": "2026-06-26T06:25:00Z", "device": "edge-1", "z_score": 4.4 }
],
"count": 2
}
}

Flow Analytics​

Pre-aggregated NetFlow analytics. These return answer shapes (ranked pairs, application breakdowns, time-buckets), not raw rows, so an agent need not do aggregation itself. Flow records are enriched with application identity (appid), geo/ASN, and transport performance (RTT, retransmits).

top_talkers​

Top bandwidth consumers.

NameTypeDefaultDescription
minutesnumber60Lookback (≤10080)
directionstring"source"source or destination
limitnumber10Results (≤1000)

top_conversations​

Source↔destination pairs ranked by volume, enriched with application, destination org/country, and performance (avg RTT in ms, retransmits).

NameTypeDefaultDescription
minutesnumber60Lookback
ipstring–Restrict to one host's conversations (address or CIDR)
limitnumber20Results (≤200)
{
"success": true,
"data": {
"conversations": [
{ "src": "192.168.100.132", "dst": "104.16.6.34", "app": "HTTPS",
"proto": "TCP", "dst_org": "Cloudflare", "dst_country": "United States",
"bytes": "2.4 GB", "bytes_raw": 2576980378, "packets": 1850000,
"flows": 12500, "avg_rtt_ms": 24.6, "retransmits": 312 }
],
"count": 1, "window_minutes": 60
}
}

top_applications​

Traffic grouped by application (appid) with share of total bytes.

NameTypeDefaultDescription
minutesnumber60Lookback
ipstring–Restrict to one host
limitnumber20Results (≤100)

Each item: app, category, bytes, flows, pct_of_total.

flow_timeseries​

Volume and transport performance bucketed over time. The primary tool for causal reasoning — it reveals the onset time of a surge and whether RTT/retransmits rose alongside volume.

NameTypeDefaultDescription
src_ipstring–Filter by source
dst_ipstring–Filter by destination
ipstring–Filter by host on either end
minutesnumber60Lookback
bucket_secondsnumber60Bucket width (60–3600)

Returns series[] of { time, bytes, packets, flows, avg_rtt_ms, retransmits }. Emits low_sample when fewer than 3 buckets.

flow_ports​

Destination-port distribution and distinct-port count for a host — scan/exfil reasoning.

NameTypeDefaultDescription
ipstringYesHost to analyze
minutesnumber60Lookback
directionstring"destination"destination = ports the host connected to; source = ports it served
limitnumber25Results (≤200)

Returns distinct_ports, a ports[] breakdown, and a high_port_fanout caveat when the distinct count exceeds 100.

flow_performance​

Aggregate transport health for a host or conversation: RTT (avg/min/max ms), retransmits and retransmit %, TCP window range — without packet capture.

NameTypeDefaultDescription
src_ip / dst_ip / ipstring–Scope to a flow or host
minutesnumber60Lookback

Emits high_retransmit (>2%), high_rtt (>150 ms), or no_data.

search_flows​

Raw 5-tuple flow search (the escape hatch when an aggregate doesn't fit).

NameTypeDefaultDescription
src_ip / dst_ipstring–Address or CIDR
src_port / dst_portnumber–0–65535
protocolstring–tcp, udp, icmp, gre, esp, ah, sctp
minutesnumber60Lookback
limitnumber100Results (≤5000)

flow_summary​

Aggregate totals over a window: bytes, packets, flows, unique sources/destinations.

NameTypeDefaultDescription
minutesnumber60Lookback

SNMP Metrics​

metric_timeseries​

Bucketed device metric trend (not just the latest point). Emits a stale caveat when the device's most recent poll is older than ~10 minutes — distinguishing "actually flat" from "not being polled."

NameTypeDefaultDescription
devicestringYesDevice name
interface_namestring–Restrict to one interface
metricstring"cpu"cpu, memory, if_in, if_out
minutesnumber60Lookback
bucket_secondsnumber300Bucket width (60–3600)

Returns series[] of { time, value }, plus last_poll and caveats.

get_interface_errors​

Interface error/discard counters for a device, only for interfaces with non-zero errors. Emits clean when none.

NameTypeDefaultDescription
devicestringYesDevice name
minutesnumber60Lookback

Each interface: in_errors, out_errors, in_discards, out_discards, oper_status, last_poll.

get_snmp_metrics​

Recent raw SNMP rows.

NameTypeDefaultDescription
devicestring–Filter by device
minutesnumber60Lookback
limitnumber100Results (≤5000)

get_interface_utilization​

Per-interface octet counters over time for a device.

NameTypeDefaultDescription
devicestringYesDevice name
interface_namestring–Specific interface
minutesnumber60Lookback

Configuration​

get_recent_config_changes​

New configuration versions across devices in a window — the "what changed" sweep, usually the first root-cause question.

NameTypeDefaultDescription
minutesnumber1440Lookback (default 24h, ≤30d)
devicestring–Filter by device
limitnumber50Results (≤200)

Each change: device, version, source, config_hash, config_bytes, changed_at.

get_device_config​

A device's configuration content at the latest or a specific version. Large configs are capped at 20 KB with a truncated caveat.

NameTypeRequiredDescription
devicestringYesDevice name
versionnumberNoSpecific version (defaults to latest)

get_config_diff​

Line-level diff between two configuration versions (defaults to the two most recent). Returns added/removed lines, not the whole file.

NameTypeRequiredDescription
devicestringYesDevice name
from_versionnumberNoOlder version (defaults to second-newest)
to_versionnumberNoNewer version (defaults to newest)

Returns added_lines, removed_lines (each capped at 200), added_count, removed_count, and version timestamps.


Logs​

search_logs​

Filter syslog by host, severity, mnemonic, or message substring.

NameTypeDefaultDescription
hoststring–Device/host name
severitystring–emergency, alert, critical, error, warning, notice, info, debug
mnemonicstring–Log mnemonic (e.g. a platform event code)
containsstring–Case-insensitive substring of the message
minutesnumber60Lookback
limitnumber100Results (≤1000)

Confirm your platform's actual severity values match the enum above; if your devices emit a different scheme, filter on mnemonic or contains instead.

get_log_bursts​

Host/mnemonic combinations with abnormally high log volume (≥10 in the window) — surfaces a device "screaming." Emits quiet when nothing reaches the threshold.

NameTypeDefaultDescription
minutesnumber60Lookback
hoststring–Restrict to one host
limitnumber25Results (≤100)

Context (Maintenance, Signals, IPAM)​

get_maintenance_windows​

Active (and optionally all) maintenance windows. An agent should check this before declaring anything down, so planned work is not reported as an outage. Indefinite windows (no end time) are flagged.

NameTypeDefaultDescription
active_onlybooleantrueOnly currently-active windows
limitnumber50Results (≤200)

Each window: scope_type (device/site/global), scope (resolved name), reason, starts_at, ends_at, indefinite, created_by. Emits indefinite_windows_present when applicable.

get_signals​

Raw anomaly signals before correlation — including signals that did not roll up into an incident. Useful for "what's firing right now" and for finding unattached anomalies.

NameTypeDefaultDescription
minutesnumber60Lookback
sourcestring–snmp, flow, syslog, synthetic, config
signal_typestring–Specific signal type
devicestring–Filter by device
unattached_onlybooleanfalseOnly signals not linked to an incident
limitnumber50Results (≤200)

Each signal: detected_at, source, type, family, severity, z_score, current_value, baseline, occurrences, device, site, interface, and incident_id/attached.

lookup_ip​

Resolve an IP to its IPAM context: address record (hostname, MAC, vendor, owner, status, linked device), the containing prefix (name, role, VRF, site), and the resolved VLAN.

NameTypeRequiredDescription
ipstringYesIP address or CIDR
{
"success": true,
"data": {
"ip": "192.168.100.221",
"address_record": { "hostname": "app-server-3", "device_name": "cat9300-sw01",
"vendor": "Dell", "status": "active" },
"containing_prefix": { "prefix": "192.168.100.0/24", "name": "HQ Servers",
"role": "server", "vrf": "Global", "site_name": "HQ" },
"vlan": { "vlan_id": 100, "name": "SERVERS" }
}
}

Synthetic Monitoring​

get_synthetic_tests​

List configured synthetic tests.

NameTypeDefaultDescription
typestring"all"icmp, http, ping, tcp, dns, or all
enabledboolean–Filter by enabled status

get_synthetic_results​

Recent synthetic test results (ping and HTTP).

NameTypeDefaultDescription
test_typestring"all"icmp/ping, http, or all
minutesnumber60Lookback
limitnumber50Results (≤1000)

Raw Queries (Admin)​

Both raw-query tools enforce read-only SQL: leading comments are stripped, only SELECT/WITH is permitted, multiple statements are rejected, and mutating keywords are blocked. Disallowed queries return 400.

query_postgres — admin​

NameTypeRequiredDescription
querystringYesA single read-only SQL statement

query_clickhouse — admin​

NameTypeRequiredDescription
querystringYesA single read-only SQL statement

Code Examples​

Python client​

import requests
from typing import Optional, Dict, Any

class ChompyClient:
def __init__(self, base_url: str, api_key: str):
self.base_url = base_url.rstrip("/")
self.headers = {
"Authorization": f"Bearer {api_key}",
"Content-Type": "application/json",
}

def call_tool(self, tool_name: str, params: Optional[Dict] = None) -> Dict[str, Any]:
resp = requests.post(
f"{self.base_url}/api/mcp/tools/{tool_name}",
headers=self.headers,
json=params or {},
)
resp.raise_for_status()
return resp.json()

# Investigation helpers
def open_incidents(self, limit=10):
return self.call_tool("get_incidents", {"state": "open", "limit": limit})

def incident_timeline(self, incident_id):
return self.call_tool("get_incident_timeline", {"incident_id": incident_id})

def conversations_for(self, ip, minutes=60):
return self.call_tool("top_conversations", {"ip": ip, "minutes": minutes})

def recent_config_changes(self, minutes=1440):
return self.call_tool("get_recent_config_changes", {"minutes": minutes})


if __name__ == "__main__":
client = ChompyClient("https://chompy.example.com", "chmp_your_api_key")

# Start from the reasoning layer
incidents = client.open_incidents()["data"]["incidents"]
for inc in incidents:
print(inc["incident_id"], inc["headline"], inc["severity"])
timeline = client.incident_timeline(inc["incident_id"])["data"]
if timeline.get("ordering_warning"):
print(" ⚠ ordering:", timeline["ordering_warning"])

JavaScript / Node.js client​

const axios = require("axios");

class ChompyClient {
constructor(baseUrl, apiKey) {
this.baseUrl = baseUrl.replace(/\/$/, "");
this.headers = {
Authorization: `Bearer ${apiKey}`,
"Content-Type": "application/json",
};
}

async callTool(toolName, params = {}) {
const { data } = await axios.post(
`${this.baseUrl}/api/mcp/tools/${toolName}`,
params,
{ headers: this.headers }
);
return data;
}

openIncidents(limit = 10) {
return this.callTool("get_incidents", { state: "open", limit });
}
conversationsFor(ip, minutes = 60) {
return this.callTool("top_conversations", { ip, minutes });
}
}

(async () => {
const client = new ChompyClient("https://chompy.example.com", "chmp_your_api_key");
const health = await client.callTool("network_health_summary");
console.log(health.data.device_status);
})().catch(console.error);

Anthropic tool-use (function calling)​

Expose Chompy tools to Claude by mapping each MCP tool to a tool definition. The model picks the tool; your handler proxies to the HTTP API.

import anthropic, requests, json

CHOMPY_URL = "https://chompy.example.com"
CHOMPY_KEY = "chmp_your_api_key"

def call_chompy(tool_name, params=None):
r = requests.post(
f"{CHOMPY_URL}/api/mcp/tools/{tool_name}",
headers={"Authorization": f"Bearer {CHOMPY_KEY}",
"Content-Type": "application/json"},
json=params or {},
)
return r.json()

tools = [
{
"name": "get_incidents",
"description": "List correlated network incidents with their root-cause signal. "
"Start here when investigating network problems.",
"input_schema": {
"type": "object",
"properties": {
"state": {"type": "string", "enum": ["open", "acknowledged", "resolved", "all"]},
"limit": {"type": "integer"},
},
},
},
{
"name": "get_incident_timeline",
"description": "Ordered signals for an incident. Check the ordering_warning field: "
"if present, the causal ordering may be inverted and the root cause "
"should not be trusted to have produced earlier signals.",
"input_schema": {
"type": "object",
"properties": {"incident_id": {"type": "string"}},
"required": ["incident_id"],
},
},
{
"name": "top_conversations",
"description": "Source↔destination pairs by volume, with application and RTT/retransmits. "
"Use after identifying a busy host to see who it talks to and whether the "
"transport is healthy.",
"input_schema": {
"type": "object",
"properties": {"ip": {"type": "string"}, "minutes": {"type": "integer"}},
},
},
{
"name": "get_maintenance_windows",
"description": "Active maintenance windows. Check before concluding a device is down — "
"planned work is expected, not an outage.",
"input_schema": {"type": "object", "properties": {"active_only": {"type": "boolean"}}},
},
]

client = anthropic.Anthropic()
SYSTEM = (
"You are a network operations assistant. Begin investigations from get_incidents. "
"Always honor caveat fields: treat 'stale' status as possibly-missing data, check "
"maintenance windows before declaring outages, and never assume causality when an "
"ordering_warning is present."
)

def run(user_message):
messages = [{"role": "user", "content": user_message}]
while True:
resp = client.messages.create(
model="claude-sonnet-4-6", max_tokens=1024,
system=SYSTEM, tools=tools, messages=messages,
)
if resp.stop_reason == "tool_use":
messages.append({"role": "assistant", "content": resp.content})
results = []
for block in resp.content:
if block.type == "tool_use":
out = call_chompy(block.name, block.input)
results.append({
"type": "tool_result",
"tool_use_id": block.id,
"content": json.dumps(out),
})
messages.append({"role": "user", "content": results})
continue
return "".join(b.text for b in resp.content if b.type == "text")

print(run("Are there any open incidents, and is the root cause trustworthy?"))

Error Handling​

HTTP status codes​

CodeMeaning
200Success
400Bad request — invalid parameter, disallowed SQL, or failed validation
401Unauthorized — missing/invalid/expired/revoked API key
403Forbidden — key lacks the required permission, or tool not in its allowed list
404Not found — unknown tool, or a referenced entity (incident, device, config) does not exist
429Rate limit exceeded
500Server error

Error response format​

{ "error": "Description of what went wrong" }

Validation failures are deterministic and safe to surface to an agent — e.g. Invalid src_ip: must be an IPv4/IPv6 address or CIDR, Only SELECT/WITH (read-only) queries are allowed, Incident not found: WO-9999-99999. Every call (including rejections) is recorded in the audit log.


Rate Limits​

  • Default: 1,000 requests per hour per API key (configurable per key).
  • The counter resets one hour after the first request in a window.
  • Exceeding the limit returns 429.

Best practices for limits​

  1. Cache slow-changing results (device lists, health) for 30–60 seconds.
  2. Prefer aggregate tools (top_conversations, flow_timeseries) over many search_flows calls.
  3. Scope time windows to what you need.
  4. Request a higher per-key limit for production agents.

Best Practices​

1. Start from the reasoning layer​

Begin investigations with get_incidents / get_incident_details / get_incident_timeline, then drill into telemetry. The platform has already correlated signals — don't re-derive that from raw flow/SNMP.

2. Always honor caveats​

A populated caveats array or ordering_warning is a signal to qualify your conclusion. In particular: treat stale as "this may be missing data, not reality"; check get_maintenance_windows before declaring an outage; and never assert causality across an ordering_warning.

3. Use the least-privileged key​

Give agents a read_only key (optionally narrowed with allowed_tools). Diagnosis needs no write access. Reserve admin for operators who need raw SELECTs.

4. Prefer aggregates over raw rows​

Pre-aggregated tools return the answer shape and cost far less context than raw-row tools. Reach for search_flows/get_snmp_metrics only when an aggregate doesn't fit.

5. Right-size time and limits​

Use short windows (5–15 min) for real-time questions and longer ones (1–24 h) for analysis. Limits are clamped server-side, but asking for less keeps responses lean.

6. Secure and rotate keys​

Never commit keys; use secrets management; rotate periodically; use separate keys per application. Review the audit log regularly.


Roadmap: Agentic Action Tiers​

This API is currently a read/reasoning surface (plus acknowledge_alert). The action surface is being built out in tiers, each gated and audited:

Tier 1 — Read & reason (available now). Everything in this document: telemetry, flow/metric analytics, config, logs, incidents, IPAM, context. An agent can fully diagnose without changing anything.

Tier 2 — Safe, reversible actions (planned). New agent_read / agent_act permission tiers (scoped via the existing allowed_tools mechanism). Tools: incident acknowledge/resolve/add_note, set_maintenance/clear_maintenance, and the active probes run_ping_test / run_traceroute / run_http_test (ported from the stdio server). Each will require a reason, support dry_run, be idempotent on retry, and be fully audited.

Tier 3 — Network-mutating actions (gated by human approval). Config push, interface state changes, and similar will not be direct agent tools. Instead a propose_action tool will write a recommendation to a human-approval queue. The agent decides what should happen; a person commits it. This keeps a reasoning error a rejected suggestion rather than an outage.


This document reflects the current mcpHttpApi.js surface: 29 tools across devices, health, alerts, incidents, flow analytics, SNMP, configuration, logs, context, IPAM, synthetics, and raw queries — with three-state device status, read-only SQL enforcement, input validation, and inline caveat fields.