Chompy MCP HTTP API Documentation
Introduction
The Chompy Model Context Protocol (MCP) HTTP API exposes your network observability platform — devices, flow, SNMP metrics, syslog, configuration history, correlated incidents, and IPAM — to AI agents, internal LLMs, and automation workflows over a simple authenticated REST surface.
This API is designed for agentic use. Beyond raw telemetry, it surfaces the platform's reasoning layer: correlated incidents with root-cause signals, ordered signal timelines, and pre-aggregated analytics (top conversations, application breakdowns, time-series). Many tools also return caveat fields — inline epistemic flags (stale, low_sample, high_retransmit, indefinite_windows_present, ordering_warning, …) that tell an agent how much to trust a result, so it reasons soundly instead of taking every number at face value.
Note on surfaces. This document describes the HTTP API (
/api/mcp/tools/...), authenticated by an MCP API key. It is distinct from the stdio MCP server used by desktop LLM clients. Active-probe tools (run_ping_test,run_traceroute,run_http_test) currently live in the stdio server and are not part of this HTTP surface; they are listed in the roadmap below.
Table of Contents
- Getting Started
- Authentication & Permissions
- API Reference
- Caveat Fields
- Tool Reference
- Code Examples
- Error Handling
- Rate Limits
- Best Practices
- Roadmap: Agentic Action Tiers
Getting Started
Base URL
https://your-chompy-instance.com/api/mcp
Quick Start
- Obtain an MCP API key from your Chompy administrator (keys are prefixed
chmp_). - Make your first request:
curl -X POST https://your-chompy-instance.com/api/mcp/tools/network_health_summary \
-H "Authorization: Bearer chmp_your_api_key" \
-H "Content-Type: application/json"
All tool calls are POST requests with a JSON body of parameters (an empty {} is fine for tools that take none).
Authentication & Permissions
All requests authenticate via a Bearer token (the MCP API key):
Authorization: Bearer chmp_your_api_key
Each key carries a permission level. Levels are hierarchical — higher levels include everything below them. A key may also be restricted to an explicit allowed_tools list, in which case it can call only those tools (within its level).
| Level | Includes | Use for |
|---|---|---|
read_only | All read tools (telemetry, incidents, analytics, config, logs, context, IPAM) | Dashboards, monitoring agents, investigation/diagnosis |
read_write | read_only + acknowledge_alert | Agents that triage and acknowledge |
admin | read_write + query_postgres, query_clickhouse | Operators needing raw SELECT access |
Scoping agents. For an autonomous agent, issue a
read_onlykey (optionally narrowed viaallowed_tools) so it can observe and diagnose the entire estate but cannot mutate anything. Action capabilities are introduced separately — see the roadmap.
Management endpoints (/keys, /audit) require an admin session token, not an MCP API key — they are operator surfaces, not agent surfaces.
API Reference
List available tools
GET /api/mcp/tools
Returns every tool, the permission it requires, and whether your key can call it. Requires a valid MCP API key.
{
"tools": [
{ "name": "get_devices", "permission_required": "read_only", "available": true },
{ "name": "query_postgres", "permission_required": "admin", "available": false }
],
"your_permission": "read_only"
}
Execute a tool
POST /api/mcp/tools/{tool_name}
Headers
Authorization: Bearer chmp_your_api_key
Content-Type: application/json
Body: tool-specific parameters (see Tool Reference).
Success response
{ "success": true, "data": { /* tool-specific */ } }
Error response
{ "error": "Description of what went wrong" }
Caveat Fields
Many tools include a caveats array (and some include named flags like ordering_warning) in their data. These are machine-readable epistemic signals: they describe limitations of the result so an agent can weigh it correctly. Treat a populated caveat as a reason to qualify conclusions, gather corroborating evidence, or widen a time window.
| Caveat | Meaning |
|---|---|
stale | The underlying data's most recent sample is older than expected (e.g. a device that has stopped being SNMP-polled). A down/flat reading may reflect missing data, not reality. |
no_data | No matching records in the window. |
low_sample | Too few data points to establish a trend. |
high_retransmit | TCP retransmission rate is elevated — degraded transport. |
high_rtt | Average round-trip time is elevated. |
high_port_fanout | A host contacted an unusually large number of distinct ports — possible scan/exfil. |
clean | Explicitly nothing wrong found (e.g. no interface errors). |
quiet | No activity reached a threshold (e.g. no log bursts). |
truncated | Output was capped; more data exists. |
indefinite_windows_present | One or more maintenance windows have no end time and will suppress incidents until manually cleared. |
ordering_warning | In an incident timeline, a symptom signal predates the designated root cause — causal ordering may be inverted. Do not assume the root cause produced earlier signals. |
Tool Reference
Conventions. Time windows are specified in
minutesand are bounded server-side.limitvalues are clamped to safe maxima. IPs accept addresses or CIDR. All string inputs are validated; invalid input returns400.
Devices & Health
get_devices
List inventory devices with live status pulled from recent SNMP polls.
| Name | Type | Default | Description |
|---|---|---|---|
device_type | string | – | Filter by type substring (router, switch, firewall) |
site | string | – | Filter by site name substring |
role | string | – | Filter by role substring |
limit | number | 50 | Max results (≤1000) |
Device status is three-state, derived from snmp_metrics.device_status:
| Value | Label | Meaning |
|---|---|---|
2 | up | Healthy — reachable by both ICMP and SNMP |
1 | degraded | Reachable by ICMP or SNMP, but not both (partial) |
0 | down | Not reachable by ICMP or SNMP |
| (no recent poll) | unknown | No SNMP data in the lookback window |
degradedis a real warning state and often precedes a full outage — do not treat it as either up or down.
Example
curl -X POST https://chompy.example.com/api/mcp/tools/get_devices \
-H "Authorization: Bearer chmp_xxx" -H "Content-Type: application/json" \
-d '{"device_type": "switch", "limit": 10}'
Response
{
"success": true,
"data": {
"devices": [
{ "id": 12, "name": "cat9300-sw01", "ip": "192.168.100.2",
"type": "switch", "role": "access", "site": "HQ",
"status": "up", "cpu": 14.2, "memory": 38.5 }
],
"count": 1
}
}
get_device_details
Detailed record for one device, including its interfaces.
| Name | Type | Required | Description |
|---|---|---|---|
device_name | string | Yes | Device name or management IP (substring match on name) |
Returns device (name, ip, type, role, site, three-state status, cpu, memory) and an interfaces array.
network_health_summary
Estate-wide situational summary. No parameters.
Returns total_devices, a device_status breakdown (using the three-state labels), active_alerts by severity, flow_stats (last hour), and synthetic_tests success rate.
{
"success": true,
"data": {
"total_devices": 24,
"device_status": [
{ "status": "up", "count": 1 },
{ "status": "degraded", "count": 1 },
{ "status": "down", "count": 20 },
{ "status": "unknown", "count": 2 }
],
"active_alerts": [ { "severity": "critical", "count": 1 } ],
"flow_stats": { "flows_last_hour": 285000, "bytes_last_hour": "45.2 GB" },
"synthetic_tests": { "successful": 142, "total": 150, "success_rate": "94.7" }
}
}
Alerts
get_alerts
Threshold-rule alert instances.
| Name | Type | Default | Description |
|---|---|---|---|
severity | string | "all" | Filter by severity, or all |
state | string | "active" | active (triggered/active), a specific state, or all |
limit | number | 25 | Max results (≤500) |
acknowledge_alert — read_write
Acknowledge an active alert.
| Name | Type | Required | Description |
|---|---|---|---|
alert_id | number | Yes | Alert instance ID |
acknowledged_by | string | No | Who/what acknowledged (defaults to API) |
notes | string | No | Acknowledgment note |
Incidents & Reasoning
These tools expose the correlated incident layer produced by the Problem Assembler — the platform's reasoning output, not raw telemetry. This is usually where an agent should start an investigation.
get_incidents
List correlated incidents (WO-YYYY-NNNNN) with their root-cause signal.
| Name | Type | Default | Description |
|---|---|---|---|
state | string | "open" | open, acknowledged, resolved, or all |
severity | string | "all" | Filter by severity |
minutes | number | – | Restrict to incidents opened in the last N minutes |
limit | number | 25 | Max results (≤200) |
Each incident includes incident_id, headline, state, severity, confidence, threat_type, signal_count, correlation_method, the root_cause_signal (type, family, detected_at), and affected device/site IDs.
get_incident_details
Full record for one incident, including AI-synthesized root cause, blast radius, remediation steps, and fault_analysis (path/topology) when present.
| Name | Type | Required | Description |
|---|---|---|---|
incident_id | string | number | Yes | The WO-… ID or numeric problem ID |
get_incident_timeline
The ordered signals that compose an incident, each with its role (root_cause / symptom), timing, severity, and key metrics. Includes the ordering_warning caveat when a symptom predates the root cause.
| Name | Type | Required | Description |
|---|---|---|---|
incident_id | string | number | Yes | The WO-… ID or numeric problem ID |
{
"success": true,
"data": {
"incident_id": "WO-2026-00469",
"ordering_warning": "Root-cause signal occurs after 1 symptom(s) in time — causal ordering may be inverted...",
"signals": [
{ "signal_id": 880, "role": "symptom", "type": "ping_loss_spike",
"detected_at": "2026-06-26T05:52:00Z", "device": "edge-1", "z_score": 6.1 },
{ "signal_id": 884, "role": "root_cause", "type": "top_talker_spike",
"detected_at": "2026-06-26T06:25:00Z", "device": "edge-1", "z_score": 4.4 }
],
"count": 2
}
}
Flow Analytics
Pre-aggregated NetFlow analytics. These return answer shapes (ranked pairs, application breakdowns, time-buckets), not raw rows, so an agent need not do aggregation itself. Flow records are enriched with application identity (appid), geo/ASN, and transport performance (RTT, retransmits).
top_talkers
Top bandwidth consumers.
| Name | Type | Default | Description |
|---|---|---|---|
minutes | number | 60 | Lookback (≤10080) |
direction | string | "source" | source or destination |
limit | number | 10 | Results (≤1000) |
top_conversations
Source↔destination pairs ranked by volume, enriched with application, destination org/country, and performance (avg RTT in ms, retransmits).
| Name | Type | Default | Description |
|---|---|---|---|
minutes | number | 60 | Lookback |
ip | string | – | Restrict to one host's conversations (address or CIDR) |
limit | number | 20 | Results (≤200) |
{
"success": true,
"data": {
"conversations": [
{ "src": "192.168.100.132", "dst": "104.16.6.34", "app": "HTTPS",
"proto": "TCP", "dst_org": "Cloudflare", "dst_country": "United States",
"bytes": "2.4 GB", "bytes_raw": 2576980378, "packets": 1850000,
"flows": 12500, "avg_rtt_ms": 24.6, "retransmits": 312 }
],
"count": 1, "window_minutes": 60
}
}
top_applications
Traffic grouped by application (appid) with share of total bytes.
| Name | Type | Default | Description |
|---|---|---|---|
minutes | number | 60 | Lookback |
ip | string | – | Restrict to one host |
limit | number | 20 | Results (≤100) |
Each item: app, category, bytes, flows, pct_of_total.
flow_timeseries
Volume and transport performance bucketed over time. The primary tool for causal reasoning — it reveals the onset time of a surge and whether RTT/retransmits rose alongside volume.
| Name | Type | Default | Description |
|---|---|---|---|
src_ip | string | – | Filter by source |
dst_ip | string | – | Filter by destination |
ip | string | – | Filter by host on either end |
minutes | number | 60 | Lookback |
bucket_seconds | number | 60 | Bucket width (60–3600) |
Returns series[] of { time, bytes, packets, flows, avg_rtt_ms, retransmits }. Emits low_sample when fewer than 3 buckets.
flow_ports
Destination-port distribution and distinct-port count for a host — scan/exfil reasoning.
| Name | Type | Default | Description |
|---|---|---|---|
ip | string | Yes | Host to analyze |
minutes | number | 60 | Lookback |
direction | string | "destination" | destination = ports the host connected to; source = ports it served |
limit | number | 25 | Results (≤200) |
Returns distinct_ports, a ports[] breakdown, and a high_port_fanout caveat when the distinct count exceeds 100.
flow_performance
Aggregate transport health for a host or conversation: RTT (avg/min/max ms), retransmits and retransmit %, TCP window range — without packet capture.
| Name | Type | Default | Description |
|---|---|---|---|
src_ip / dst_ip / ip | string | – | Scope to a flow or host |
minutes | number | 60 | Lookback |
Emits high_retransmit (>2%), high_rtt (>150 ms), or no_data.
search_flows
Raw 5-tuple flow search (the escape hatch when an aggregate doesn't fit).
| Name | Type | Default | Description |
|---|---|---|---|
src_ip / dst_ip | string | – | Address or CIDR |
src_port / dst_port | number | – | 0–65535 |
protocol | string | – | tcp, udp, icmp, gre, esp, ah, sctp |
minutes | number | 60 | Lookback |
limit | number | 100 | Results (≤5000) |
flow_summary
Aggregate totals over a window: bytes, packets, flows, unique sources/destinations.
| Name | Type | Default | Description |
|---|---|---|---|
minutes | number | 60 | Lookback |
SNMP Metrics
metric_timeseries
Bucketed device metric trend (not just the latest point). Emits a stale caveat when the device's most recent poll is older than ~10 minutes — distinguishing "actually flat" from "not being polled."
| Name | Type | Default | Description |
|---|---|---|---|
device | string | Yes | Device name |
interface_name | string | – | Restrict to one interface |
metric | string | "cpu" | cpu, memory, if_in, if_out |
minutes | number | 60 | Lookback |
bucket_seconds | number | 300 | Bucket width (60–3600) |
Returns series[] of { time, value }, plus last_poll and caveats.
get_interface_errors
Interface error/discard counters for a device, only for interfaces with non-zero errors. Emits clean when none.
| Name | Type | Default | Description |
|---|---|---|---|
device | string | Yes | Device name |
minutes | number | 60 | Lookback |
Each interface: in_errors, out_errors, in_discards, out_discards, oper_status, last_poll.
get_snmp_metrics
Recent raw SNMP rows.
| Name | Type | Default | Description |
|---|---|---|---|
device | string | – | Filter by device |
minutes | number | 60 | Lookback |
limit | number | 100 | Results (≤5000) |
get_interface_utilization
Per-interface octet counters over time for a device.
| Name | Type | Default | Description |
|---|---|---|---|
device | string | Yes | Device name |
interface_name | string | – | Specific interface |
minutes | number | 60 | Lookback |
Configuration
get_recent_config_changes
New configuration versions across devices in a window — the "what changed" sweep, usually the first root-cause question.
| Name | Type | Default | Description |
|---|---|---|---|
minutes | number | 1440 | Lookback (default 24h, ≤30d) |
device | string | – | Filter by device |
limit | number | 50 | Results (≤200) |
Each change: device, version, source, config_hash, config_bytes, changed_at.
get_device_config
A device's configuration content at the latest or a specific version. Large configs are capped at 20 KB with a truncated caveat.
| Name | Type | Required | Description |
|---|---|---|---|
device | string | Yes | Device name |
version | number | No | Specific version (defaults to latest) |
get_config_diff
Line-level diff between two configuration versions (defaults to the two most recent). Returns added/removed lines, not the whole file.
| Name | Type | Required | Description |
|---|---|---|---|
device | string | Yes | Device name |
from_version | number | No | Older version (defaults to second-newest) |
to_version | number | No | Newer version (defaults to newest) |
Returns added_lines, removed_lines (each capped at 200), added_count, removed_count, and version timestamps.
Logs
search_logs
Filter syslog by host, severity, mnemonic, or message substring.
| Name | Type | Default | Description |
|---|---|---|---|
host | string | – | Device/host name |
severity | string | – | emergency, alert, critical, error, warning, notice, info, debug |
mnemonic | string | – | Log mnemonic (e.g. a platform event code) |
contains | string | – | Case-insensitive substring of the message |
minutes | number | 60 | Lookback |
limit | number | 100 | Results (≤1000) |
Confirm your platform's actual
severityvalues match the enum above; if your devices emit a different scheme, filter onmnemonicorcontainsinstead.
get_log_bursts
Host/mnemonic combinations with abnormally high log volume (≥10 in the window) — surfaces a device "screaming." Emits quiet when nothing reaches the threshold.
| Name | Type | Default | Description |
|---|---|---|---|
minutes | number | 60 | Lookback |
host | string | – | Restrict to one host |
limit | number | 25 | Results (≤100) |
Context (Maintenance, Signals, IPAM)
get_maintenance_windows
Active (and optionally all) maintenance windows. An agent should check this before declaring anything down, so planned work is not reported as an outage. Indefinite windows (no end time) are flagged.
| Name | Type | Default | Description |
|---|---|---|---|
active_only | boolean | true | Only currently-active windows |
limit | number | 50 | Results (≤200) |
Each window: scope_type (device/site/global), scope (resolved name), reason, starts_at, ends_at, indefinite, created_by. Emits indefinite_windows_present when applicable.
get_signals
Raw anomaly signals before correlation — including signals that did not roll up into an incident. Useful for "what's firing right now" and for finding unattached anomalies.
| Name | Type | Default | Description |
|---|---|---|---|
minutes | number | 60 | Lookback |
source | string | – | snmp, flow, syslog, synthetic, config |
signal_type | string | – | Specific signal type |
device | string | – | Filter by device |
unattached_only | boolean | false | Only signals not linked to an incident |
limit | number | 50 | Results (≤200) |
Each signal: detected_at, source, type, family, severity, z_score, current_value, baseline, occurrences, device, site, interface, and incident_id/attached.
lookup_ip
Resolve an IP to its IPAM context: address record (hostname, MAC, vendor, owner, status, linked device), the containing prefix (name, role, VRF, site), and the resolved VLAN.
| Name | Type | Required | Description |
|---|---|---|---|
ip | string | Yes | IP address or CIDR |
{
"success": true,
"data": {
"ip": "192.168.100.221",
"address_record": { "hostname": "app-server-3", "device_name": "cat9300-sw01",
"vendor": "Dell", "status": "active" },
"containing_prefix": { "prefix": "192.168.100.0/24", "name": "HQ Servers",
"role": "server", "vrf": "Global", "site_name": "HQ" },
"vlan": { "vlan_id": 100, "name": "SERVERS" }
}
}
Synthetic Monitoring
get_synthetic_tests
List configured synthetic tests.
| Name | Type | Default | Description |
|---|---|---|---|
type | string | "all" | icmp, http, ping, tcp, dns, or all |
enabled | boolean | – | Filter by enabled status |
get_synthetic_results
Recent synthetic test results (ping and HTTP).
| Name | Type | Default | Description |
|---|---|---|---|
test_type | string | "all" | icmp/ping, http, or all |
minutes | number | 60 | Lookback |
limit | number | 50 | Results (≤1000) |
Raw Queries (Admin)
Both raw-query tools enforce read-only SQL: leading comments are stripped, only SELECT/WITH is permitted, multiple statements are rejected, and mutating keywords are blocked. Disallowed queries return 400.
query_postgres — admin
| Name | Type | Required | Description |
|---|---|---|---|
query | string | Yes | A single read-only SQL statement |
query_clickhouse — admin
| Name | Type | Required | Description |
|---|---|---|---|
query | string | Yes | A single read-only SQL statement |
Code Examples
Python client
import requests
from typing import Optional, Dict, Any
class ChompyClient:
def __init__(self, base_url: str, api_key: str):
self.base_url = base_url.rstrip("/")
self.headers = {
"Authorization": f"Bearer {api_key}",
"Content-Type": "application/json",
}
def call_tool(self, tool_name: str, params: Optional[Dict] = None) -> Dict[str, Any]:
resp = requests.post(
f"{self.base_url}/api/mcp/tools/{tool_name}",
headers=self.headers,
json=params or {},
)
resp.raise_for_status()
return resp.json()
# Investigation helpers
def open_incidents(self, limit=10):
return self.call_tool("get_incidents", {"state": "open", "limit": limit})
def incident_timeline(self, incident_id):
return self.call_tool("get_incident_timeline", {"incident_id": incident_id})
def conversations_for(self, ip, minutes=60):
return self.call_tool("top_conversations", {"ip": ip, "minutes": minutes})
def recent_config_changes(self, minutes=1440):
return self.call_tool("get_recent_config_changes", {"minutes": minutes})
if __name__ == "__main__":
client = ChompyClient("https://chompy.example.com", "chmp_your_api_key")
# Start from the reasoning layer
incidents = client.open_incidents()["data"]["incidents"]
for inc in incidents:
print(inc["incident_id"], inc["headline"], inc["severity"])
timeline = client.incident_timeline(inc["incident_id"])["data"]
if timeline.get("ordering_warning"):
print(" ⚠ ordering:", timeline["ordering_warning"])
JavaScript / Node.js client
const axios = require("axios");
class ChompyClient {
constructor(baseUrl, apiKey) {
this.baseUrl = baseUrl.replace(/\/$/, "");
this.headers = {
Authorization: `Bearer ${apiKey}`,
"Content-Type": "application/json",
};
}
async callTool(toolName, params = {}) {
const { data } = await axios.post(
`${this.baseUrl}/api/mcp/tools/${toolName}`,
params,
{ headers: this.headers }
);
return data;
}
openIncidents(limit = 10) {
return this.callTool("get_incidents", { state: "open", limit });
}
conversationsFor(ip, minutes = 60) {
return this.callTool("top_conversations", { ip, minutes });
}
}
(async () => {
const client = new ChompyClient("https://chompy.example.com", "chmp_your_api_key");
const health = await client.callTool("network_health_summary");
console.log(health.data.device_status);
})().catch(console.error);
Anthropic tool-use (function calling)
Expose Chompy tools to Claude by mapping each MCP tool to a tool definition. The model picks the tool; your handler proxies to the HTTP API.
import anthropic, requests, json
CHOMPY_URL = "https://chompy.example.com"
CHOMPY_KEY = "chmp_your_api_key"
def call_chompy(tool_name, params=None):
r = requests.post(
f"{CHOMPY_URL}/api/mcp/tools/{tool_name}",
headers={"Authorization": f"Bearer {CHOMPY_KEY}",
"Content-Type": "application/json"},
json=params or {},
)
return r.json()
tools = [
{
"name": "get_incidents",
"description": "List correlated network incidents with their root-cause signal. "
"Start here when investigating network problems.",
"input_schema": {
"type": "object",
"properties": {
"state": {"type": "string", "enum": ["open", "acknowledged", "resolved", "all"]},
"limit": {"type": "integer"},
},
},
},
{
"name": "get_incident_timeline",
"description": "Ordered signals for an incident. Check the ordering_warning field: "
"if present, the causal ordering may be inverted and the root cause "
"should not be trusted to have produced earlier signals.",
"input_schema": {
"type": "object",
"properties": {"incident_id": {"type": "string"}},
"required": ["incident_id"],
},
},
{
"name": "top_conversations",
"description": "Source↔destination pairs by volume, with application and RTT/retransmits. "
"Use after identifying a busy host to see who it talks to and whether the "
"transport is healthy.",
"input_schema": {
"type": "object",
"properties": {"ip": {"type": "string"}, "minutes": {"type": "integer"}},
},
},
{
"name": "get_maintenance_windows",
"description": "Active maintenance windows. Check before concluding a device is down — "
"planned work is expected, not an outage.",
"input_schema": {"type": "object", "properties": {"active_only": {"type": "boolean"}}},
},
]
client = anthropic.Anthropic()
SYSTEM = (
"You are a network operations assistant. Begin investigations from get_incidents. "
"Always honor caveat fields: treat 'stale' status as possibly-missing data, check "
"maintenance windows before declaring outages, and never assume causality when an "
"ordering_warning is present."
)
def run(user_message):
messages = [{"role": "user", "content": user_message}]
while True:
resp = client.messages.create(
model="claude-sonnet-4-6", max_tokens=1024,
system=SYSTEM, tools=tools, messages=messages,
)
if resp.stop_reason == "tool_use":
messages.append({"role": "assistant", "content": resp.content})
results = []
for block in resp.content:
if block.type == "tool_use":
out = call_chompy(block.name, block.input)
results.append({
"type": "tool_result",
"tool_use_id": block.id,
"content": json.dumps(out),
})
messages.append({"role": "user", "content": results})
continue
return "".join(b.text for b in resp.content if b.type == "text")
print(run("Are there any open incidents, and is the root cause trustworthy?"))
Error Handling
HTTP status codes
| Code | Meaning |
|---|---|
200 | Success |
400 | Bad request — invalid parameter, disallowed SQL, or failed validation |
401 | Unauthorized — missing/invalid/expired/revoked API key |
403 | Forbidden — key lacks the required permission, or tool not in its allowed list |
404 | Not found — unknown tool, or a referenced entity (incident, device, config) does not exist |
429 | Rate limit exceeded |
500 | Server error |
Error response format
{ "error": "Description of what went wrong" }
Validation failures are deterministic and safe to surface to an agent — e.g. Invalid src_ip: must be an IPv4/IPv6 address or CIDR, Only SELECT/WITH (read-only) queries are allowed, Incident not found: WO-9999-99999. Every call (including rejections) is recorded in the audit log.
Rate Limits
- Default: 1,000 requests per hour per API key (configurable per key).
- The counter resets one hour after the first request in a window.
- Exceeding the limit returns
429.
Best practices for limits
- Cache slow-changing results (device lists, health) for 30–60 seconds.
- Prefer aggregate tools (
top_conversations,flow_timeseries) over manysearch_flowscalls. - Scope time windows to what you need.
- Request a higher per-key limit for production agents.
Best Practices
1. Start from the reasoning layer
Begin investigations with get_incidents / get_incident_details / get_incident_timeline, then drill into telemetry. The platform has already correlated signals — don't re-derive that from raw flow/SNMP.
2. Always honor caveats
A populated caveats array or ordering_warning is a signal to qualify your conclusion. In particular: treat stale as "this may be missing data, not reality"; check get_maintenance_windows before declaring an outage; and never assert causality across an ordering_warning.
3. Use the least-privileged key
Give agents a read_only key (optionally narrowed with allowed_tools). Diagnosis needs no write access. Reserve admin for operators who need raw SELECTs.
4. Prefer aggregates over raw rows
Pre-aggregated tools return the answer shape and cost far less context than raw-row tools. Reach for search_flows/get_snmp_metrics only when an aggregate doesn't fit.
5. Right-size time and limits
Use short windows (5–15 min) for real-time questions and longer ones (1–24 h) for analysis. Limits are clamped server-side, but asking for less keeps responses lean.
6. Secure and rotate keys
Never commit keys; use secrets management; rotate periodically; use separate keys per application. Review the audit log regularly.
Roadmap: Agentic Action Tiers
This API is currently a read/reasoning surface (plus acknowledge_alert). The action surface is being built out in tiers, each gated and audited:
Tier 1 — Read & reason (available now). Everything in this document: telemetry, flow/metric analytics, config, logs, incidents, IPAM, context. An agent can fully diagnose without changing anything.
Tier 2 — Safe, reversible actions (planned). New agent_read / agent_act permission tiers (scoped via the existing allowed_tools mechanism). Tools: incident acknowledge/resolve/add_note, set_maintenance/clear_maintenance, and the active probes run_ping_test / run_traceroute / run_http_test (ported from the stdio server). Each will require a reason, support dry_run, be idempotent on retry, and be fully audited.
Tier 3 — Network-mutating actions (gated by human approval). Config push, interface state changes, and similar will not be direct agent tools. Instead a propose_action tool will write a recommendation to a human-approval queue. The agent decides what should happen; a person commits it. This keeps a reasoning error a rejected suggestion rather than an outage.
This document reflects the current mcpHttpApi.js surface: 29 tools across devices, health, alerts, incidents, flow analytics, SNMP, configuration, logs, context, IPAM, synthetics, and raw queries — with three-state device status, read-only SQL enforcement, input validation, and inline caveat fields.