Synter’s Ad MCP Benchmark 2026 opens with a rule worth pinning to the wall: “transport success, terminal success, and semantic success are different measurements.”
It is the most rigorous document anyone has published about advertising MCP servers. A frozen corpus. A provider-by-provider evidence table. Two-sided 95% Wilson intervals. A public harness, a dated methodology, and — rarest of all — a limitations section that states plainly what the study did not measure. Joel Horwitz, Synter’s founder and CEO, signs it, and the report reviews Synter alongside Google, Meta, TikTok, Amazon, Microsoft, Snap, Pipeboard, AdsMCP, and Adspirer.
The same company also documents its own record layer on a Security & Governance page — one that has been live since at least January, and that carried the identical audit-log promise in a snapshot from May 12, 2026. Four months before it benchmarked everyone else’s MCP servers, the page said this: “Audit logs are immutable and retained for 2 years.”
Both artifacts are real, and both are good work. Read together, they are the clearest picture yet of the thing this category has still not defined.
The benchmark measures reliability to two decimal places. No column in it measures the record a regulator asks for. And the one number that decides whether your record survives an audit is published three times, three different ways — on the same site, on the same day.
What the benchmark actually graded
The report’s provider table has five columns: documented surface, writes, safety posture, evidence class, and comparable production outcome. Before the conclusions, read the constraints — because they are unusually well stated:
| Dimension | What the report measured |
|---|---|
| Transport success | A connection or session stayed usable |
| Terminal success | A tool returned without an isError or an audited terminal failure |
| Semantic success | A task-specific hidden verifier confirmed the intended result and, for writes, the independent post-state |
| Safety posture | One cell per server, e.g. “Approval and platform-specific safety gates” |
| Evidence class | Source availability — Apache-2.0, MIT, BSL 1.1, or “source not public” |
Synter published its own production numbers: 60,880 successful calls out of 63,231 over a trailing 30 days — a 96.28% terminal success rate, 81.32% across “substantive” calls — and then wrote the sentence most vendors would have deleted:
“These are audited terminal outcomes. They do not establish that an ad report contained the right numbers, that a campaign mutation reached the intended state, or that a business objective was achieved. The semantic-success rate is therefore not measured.”
That is a benchmark telling you its own headline number is not a business outcome. It also reports “policy or approval” as a failure family — 15 calls, 0.6% — which is the telling detail. In a reliability framework, an approval gate is a thing that makes a call fail. In a compliance framework, an approval gate is the thing that makes a change defensible. Same mechanism, opposite sign.
The measurement a benchmark cannot make
Here is the report’s own description of the data it published:
“The committed aggregate contains no customer identifiers, request/response bodies, secrets, account IDs, or payment information.”
Read that as an engineer and it is exactly right. A benchmark has to be reproducible, shareable, and non-identifying — that is what makes it a benchmark instead of a data leak. Strip the account IDs, publish the method, let strangers rerun it.
Now read it as the person who has to answer a regulator. Every field the benchmark deliberately removed is a field a compliance record cannot exist without: which client’s account, what the request and response actually said, who authorized the change, what the value was before and after. A compliance record containing no customer identifiers is not a stronger record. It is not a record.
This is not a knock on the benchmark. It is the boundary of the form. Reproducibility and accountability pull in opposite directions, and the category has now published a first-rate artifact for one of them while publishing nothing for the other. When the standard for reliability is “can a stranger reproduce my number,” the supply side of trust is solved. The agency’s question is the other side: when the regulator asks what happened in this client’s account in March, what do you hand over?
Four months of governance documentation
Credit first, because it is earned — and because it is older than the benchmark that prompted this post.
Synter’s Security & Governance page is the most complete governance documentation any competitor in this market has published. It describes Admin / Editor / Viewer roles, per-platform approval permissions, and two operating modes: autopilot inside guardrails, or human approval required before launch and edit. Budget caps are enforced before the platform API call, “regardless of instructions.” Every action lands in a change journal with timestamp, actor, entity, field, old and new values, the model’s stated rationale, and the expected metrics delta. There is one-click rollback for the last ten changes per entity, PII redaction before data reaches a frontier model, a Trust Center, a published sub-processor list, a DPA with Standard Contractual Clauses, and SOC 2 Type II marked “in progress.”
The recent addition is Agent-Shield, an MIT-licensed integrity framework that signs every skill and playbook into HMAC-SHA256 manifests and validates signatures at preflight, before any tool call runs. The repository was created on August 5, 2026.
This maps onto the six properties we laid out for a real audit trail in August. So run the score — that is what the framework is for.
| Property | Where Synter's own documentation stands |
|---|---|
| Actor attribution | Documented — “Actor: Agent ID or user email” |
| Immutability | Claimed — “Audit logs are immutable.” No mechanism described: no hash chain, no append-only proof, no third-party attestation |
| Before/after values | Documented — “Old/New Values: Before and after snapshots” |
| Independent storage | Not addressed — the record lives inside the Synter workspace; export is on demand |
| Jurisdiction awareness | Not addressed — see below |
| Retention and export | Three published numbers that do not agree |
Two of six are unaddressed and one is a claim without a mechanism. Four have moved from a marketing sentence into documentation — and they did that before this benchmark existed, which is the part worth sitting with. Even the vendor with the most mature governance surface in the category has not specified the record in a way a regulated agency can hand to anyone.
Three retention numbers, no reconciliation
All three were live on syntermedia.ai on the same day:
| Surface | Object described | What it says |
|---|---|---|
| /security-governance | “Audit logs” | “immutable and retained for 2 years” |
| /pricing | “Audit Logs & Governance” plan row | 7-day history (Solo) · 90-day history (Scale) · Unlimited / SIEM Export (Custom) |
| /security-governance | “API request/response logs” | “Choose log retention period (0, 30, or 90 days)” |
The third row is a different object — request logs, not the audit journal — and it is fine for it to have its own window. The first two are not different objects. Both describe the audit record, one says two years, the other sells seven days on the entry plan, and neither page tells you which one governs.
Maybe the pricing row means the in-product history view while the security page means underlying storage. Maybe the reverse. No page says. So a regulated agency trying to answer “how long do you keep my record” has to reconcile three published numbers itself — and the only surface most buyers read before signing is the pricing table, where the answer is seven days.
We have been here before with this exact row. In August, “Audit Logs & Governance” with its 7-day / 90-day / Unlimited-SIEM ladder was the one compliance claim that survived a homepage redesign after the marketing language was scrubbed. It survived that redesign. And it has contradicted a standing claim on the security page, in public, for four months.
That is the finding. Not that the words are borrowed — that two published numbers describing one object have sat on the same domain since May without anyone reconciling them, while the company was building a benchmark rigorous enough to publish Wilson intervals for tool-call latency. The rigor is real. It has not been pointed at the record.
The claim is on its third scope
The vocabulary has had an unusual month. Chart it:
| Date | Surface | The claim |
|---|---|---|
| Aug 22 | Homepage + MCP page | “Spend limits, approval gates, and an audit trail on every write” |
| Sep 1 | Homepage | Re-scoped to “an audit trail — on every major platform” |
| Sep 7 | MCP page | Approval language demoted to an FAQ line; no audit-trail claim |
| Sep 12 | Homepage meta description | “spend limits, approvals, and an audit trail on every action” |
| Sep 12 | MCP page hero | “One unified MCP server with strict human spend approval gates” — the FAQ line is gone |
Two things are worth noticing. The per-action promise came back stronger than the version that was cut — “on every action” is a per-action claim, where September’s re-scope had quietly softened it to a coverage claim. And on the MCP page, the approval language climbed from an FAQ into the hero in five days.
To be fair: if the product now backs that, the words are simply accurate, and four months of security-page documentation suggests a real attempt. But a claim that has occupied three different scopes on three different surfaces in three weeks has not settled into a specification yet. It is still being tested on audiences. Get the retention number in the contract, not the hero.
Where every surface goes quiet: jurisdiction
We scanned 49,799 characters of text across the six surfaces that define this product — homepage, MCP page, pricing, Trust Center, Security & Governance, and the benchmark itself.
The words “jurisdiction,” “regulator,” “regulated,” “FCA,” “MiCA,” and “ESMA” appear zero times. “Compliance” appears six times, and every one of them is a certification or a review process: SOC 2, GDPR, CCPA, and the phrase “for compliance reviews.”
That gap is not a Synter problem; it is the category’s. Data residency — US infrastructure, EU region “on the roadmap” — is not campaign jurisdiction. The same budget edit means different things under an FCA financial promotion, a MiCA marketing communication, and an Alberta iGaming licence. The FCA said as much in June: “Accountability for regulated activities and outcomes must remain clear.” Its own August research found 44% of investors wrongly believe AI-generated financial information is regulated and 38% would invest on AI output alone. The IAB’s AAMP 2.3 called autonomy that respects a “hard approval boundary” a “critical requirement for regulated advertisers.”
A record that cannot say which rules applied to a change cannot answer any of them. That is the fourth measurement — and no benchmark column and no security page currently scores it.
Five questions before you let an agent write to a client account
Run these against any ad MCP vendor, including the one in this post and including Ott:
- Which success did you measure? If a vendor quotes a percentage, ask whether it is transport, terminal, or semantic — and whether they published a limitations section like this benchmark did. A terminal success rate is not a business outcome.
- Who does the record name? “An agent made a change” is not attribution. A named human, a named tool, a timestamp, a client.
- Where does the record live, and does it survive a platform restriction? If it lives only inside the vendor’s dashboard, ask what happens to it when Meta restricts the ad account it describes.
- How long is it kept, and which page is authoritative? Get one number, in the contract. Not three surfaces with three answers.
- Which jurisdiction’s rules governed the change, and can you export it in time? An FCA broker, a MiCA exchange, and an Alberta operator are three different rule sets on one agency’s client list.
Ott’s lane
Ott exists on the other side of this boundary. Telegram conversion tracking attributes every join and deposit to the ad that drove it, through a CAPI postback that survives the chat. Activity Logging records every campaign change — pauses, budget edits, creative swaps, bid and review changes — timestamped, attributed to the person or the AI agent by name, with before-and-after values, stored independently so the record outlives an account restriction, and exportable in about 90 seconds when someone asks. Budget Ledger and named approval gates put a name and a timestamp on sign-off instead of a chat thread. Agency Hierarchy jurisdiction tags know the difference between an FCA-authorised broker and a MiCA-licensed exchange. Flat pricing, $29–$199 a month, no per-client fees — and no retention tier to reconcile.
The honest scope note: we have not published a reproducible MCP reliability harness, and Synter’s is a genuinely good one. This argument is narrower, and it is about category definitions. A benchmark can tell you whether a tool returned. It cannot tell you who approved the change, which client it hit, which rules applied, or whether you can produce the file in time. Those four questions carry the licence, and as of this week nobody in the advertising MCP category has published an artifact that measures any of them.
The Ad MCP Benchmark scored the calls three ways. Your regulator asks for the record.
Try Ott free · See how Activity Logging records every campaign change