Skill Evaluation
Run one skill standalone on a payload and see what it would do: POST /v1/skills/{skill_id}/evaluate (request body, response envelope, errors, limits by plan) and GET /v1/skills/evaluate/examples. What the playground calls.
Base URL: https://engine.acefleet.dev
POST /v1/skills/{skill_id}/evaluate
Auth: dev key (Authorization: Bearer ace_dev_…).
Run ONE skill standalone on a payload and see what it would do — free, on any dev key tier (sandbox and over-budget keys included), and it never calls a provider. skill_id is one of pii_ner, prompt_compaction, agent_trajectory_compaction, output_budget, reasoning_effort, llm_router, injection_guard, multi_agent_guard, agent_trajectory_router, geo_fence_compliance, and read-only semantic_cache and skill_knowledge_graph. The body is lenient: {system?, messages[], tools?, model?, params?, max_tokens?, max_completion_tokens?, reasoning_effort?, thinking?}, or an Anthropic, OpenAI Responses, Gemini or Bedrock Converse body as sent; anything it has to guess at (a bare array, a string, a role alias, an unknown field) is accepted and named in warnings, never rejected. The answer is one envelope for every skill: verdict (changed | unchanged | flagged | skipped), a one-sentence summary, a stable reason code, metrics (tokens_before, tokens_after, latency_ms), changes[] (where, before, after, why), the skill's own decision, output (the rewritten body in your input's shape, only when changed), warnings, request_id (ev_…), docs_url and note. Standalone means simplified: no other skill runs, there is no session state or cache, and params are the defaults unless you pass params — production results can differ. Errors: 401 invalid_developer_key (missing or invalid key; the body says how to get a free one), 404 for an unknown skill (the body lists the valid ids), 413 over 2 MB, 429 rate_limited on the free plan only. Limits by plan: Free (the Evaluation and Developer tiers) — 100 evaluations a minute per key on the hosted gateway; Pro / Team and Enterprise — unlimited; an enterprise (on-prem) deployment imposes no limit at all, whatever the plan. The free-plan 429 carries Retry-After and an upgrade call to action, as fields both inside error and at the top level: plan (free), limit_per_minute, upgrade_url (https://acefleet.dev/pricing) and retry_after_s, with error.type rate_limited and an error.message naming the limit and the upgrade. The plan is read from your organization's subscription and re-checked when a key is refused, so an upgrade lifts the limit within a minute. Everything else is a 200: a skill that runs past its time budget or fails internally answers verdict: skipped with reason timeout or internal_error. Each run writes one metadata-only row (no text) kept 72 hours, listed by GET /api/v1/requests?source=skill_evaluation.
GET /v1/skills/evaluate/examples
Auth: none.
Example payloads for the evaluate endpoint, per skill: {title, request, expected: {verdict, reason}}. Static and public — the same fixtures the gateway's contract tests pin, and what the playground's example menu loads.
Try it in the browser at https://acefleet.dev/playground.