Ledgent
Ledgent is a secure agentic integration layer that exposes Salesforce and billing operations as MCP tools. It enforces action-level authorization, idempotency, PII tokenization, and full audit logging, enabling AI agents to safely perform writes with human approval for high-risk actions.
README
Ledgent
Secure Agentic Integration Layer for Salesforce + Enterprise Systems
<!-- Replace USERNAME above with your GitHub username/org once this is pushed. -->
Target role: Application Engineer, General Software Development — Google Corp Eng Cost: $0, zero credit card required anywhere in the stack.
The Problem
AI agents are moving from answering questions to taking actions — updating a Salesforce record, issuing a refund, changing an entitlement, calling another system. The moment an agent can act, three things become non-negotiable:
- Every action must be authorized — not just "the agent has an API key," but "this specific action, on this specific record, is permitted for this agent."
- Every action must be safe to retry — agent loops, network failures, and duplicate webhooks must never cause the same refund or cancellation to fire twice.
- Every action must be explainable and reversible — a human reviewer needs to see what the agent did, why, and on what evidence, after the fact.
Ledgent sits between an AI agent and Salesforce (plus one real third-party billing system) and enforces those three guarantees before any write reaches a live system.
Design Goals
- No write action executes without passing authorization + validation.
- Every decision is explainable: what, why, based on what evidence.
- PII is tokenized before it reaches the LLM — never sent raw.
- The same action cannot execute twice (idempotency, not best-effort dedup).
- Every action is fully auditable and reconstructable after the fact.
- Failures land in a known state (retryable / quarantined / needs-human) — never a silent crash or silent drop.
- High-risk actions require explicit human approval before execution.
What This Is Not
- Not a chatbot — the LLM proposes actions, it does not have direct write access.
- Not an IAM replacement — auth is necessary but not sufficient; this adds action-level policy on top.
- Not a distributed-systems research project — no claimed "business-state serializability" or "counterfactual simulation."
- Not Kafka or Kubernetes for keyword padding — a Postgres-backed audit log and a single Render web service are sufficient at this scale.
- Not fake scale — if the load test runs 5,000 events, the README says
5,000 (see
docs/BENCHMARKS.md— no load test has been run, and that page says so plainly rather than inventing a number).
Architecture
USER / TRIGGER
│
▼
API GATEWAY (FastAPI)
OAuth · RBAC · rate limit
│
▼
AGENT ORCHESTRATOR
│
┌──────────┼──────────┐
▼ ▼ ▼
MCP TOOLS LLM (cloud TASK QUEUE
│ + Ollama (ARQ/Redis)
│ fallback)
▼
Salesforce / Billing service (billing_service/)
│
▼
POLICY & APPROVAL GATE
RBAC · HMAC-verified approval · idempotency (Redis SET NX EX)
│
▼
SECURE EXECUTION
│
▼
OUTCOME VERIFIER → AUDIT / EVENT LOG (Postgres)
Repo Layout
ledgent/
├── app/ # Main FastAPI application (Ledgent)
│ ├── main.py
│ ├── config.py
│ └── integrations/
│ └── salesforce.py
├── billing_service/ # Separate service simulating a real third-party system
│ ├── main.py
│ ├── models.py
│ └── store.py
├── tests/
├── adr/ # Architecture Decision Records — see index below
├── docker/ # Dockerfiles (app + billing) and local-dev compose
│ ├── Dockerfile
│ ├── Dockerfile.billing
│ ├── requirements-billing.txt
│ └── docker-compose.yml # local Postgres + Redis for development
├── docs/
│ ├── DEPLOYMENT.md # Render + CI/CD setup walkthrough
│ ├── WINDOWS.md # PowerShell + local-dev notes for Windows
│ ├── BENCHMARKS.md # every measured number, with the command to reproduce it
│ └── DEMO_SCRIPT.md # 60-90s demo video shot list
├── render.yaml # Render Blueprint (both services, DB, Redis)
├── Makefile # local reproduction of every CI gate
└── .github/workflows/
└── ci-cd.yml # lint -> typecheck -> security -> tests -> build -> deploy
Status
Phase 0 — Foundations: done (this README, repo structure, ADRs, local Postgres/Redis via docker-compose).
Phase 1 — Real Integrations: done (Salesforce OAuth2 client via a real Developer Edition Connected App; a real billing service with genuine state transitions on subscriptions and refunds).
Phase 2 — MCP Tool Layer: done (Salesforce + billing operations exposed
as MCP tools with structured risk_level / required_scope / PII metadata
that the Phase 4 policy gate reads).
Phase 3 — Agent + Execution State Machine: done (explicit state machine, every transition audited).
Phase 4 — Policy Gate + RBAC: done (RBAC scope check, then risk-based
approval gate; scope violations are rejected before ever reaching the risk
check — see adr/0008-rbac-policy-gate.md).
Phase 5 — Security: HMAC, Idempotency, PII Tokenization: done (atomic
Redis SET NX EX idempotency, HMAC-SHA256-verified approval webhooks,
regex-based PII tokenization before anything reaches the reasoning step).
Phase 6 — Async Execution + Audit: done.
POST /v1/agent/casesnow enqueues a real ARQ job and returns202 Acceptedimmediately; a separate worker process (arq app.queue.worker_settings.WorkerSettings) runs the case and the client pollsGET /v1/agent/cases/{action_id}.POST /v1/agent/cases/synckeeps the old blocking behavior for quick local testing. Seeadr/0012-arq-async-execution.md.- Every state transition, policy decision, approval, and execution
outcome is written to an append-only Postgres/SQLite audit log,
reconstructable in order via
GET /v1/audit/{action_id}. Seeadr/0013-postgres-audit-log.md. GET /metricsexposes Prometheus-format counters and one histogram (policy decisions, state transitions, idempotency blocks, case duration, queue throughput). Seeadr/0014-prometheus-metrics-scope.md.- Structured JSON logging via
app/observability/logging_config.py.
See adr/ for the reasoning behind each infrastructure choice, and the
project plan for the full 10-phase roadmap.
Phase 7 — Testing: done. Five separate test tiers — see
adr/0015-testing-strategy.md:
- Unit (pre-existing, Phases 0-6): 64 tests, every external dependency mocked/faked.
- Contract (
tests/test_mcp_contracts.py, 8 tests): every MCP tool now carries a real, enforcedinput_schema/output_schema(app/mcp/schemas.py) — a gap between the plan and the code that existed through Phases 0-6 (seeadr/0017-mcp-tool-contract-schemas.md). A malformed call fails fast withToolContractErrorbefore ever reaching Salesforce or the billing service. - Adversarial (
tests/test_adversarial.py, 5 tests): the three cases the plan names by name — out-of-scope write, retried action, forged/replayed approval webhook — run end to end through the real orchestrator/API rather than in isolation. - Concurrency (
tests/test_concurrency.py, 4 tests): realThreadPoolExecutor-driven concurrent calls, not sequential double-calls standing in for it. Writing this tier surfaced a real bug — a state-machine race independent of the Redis idempotency guarantee — found, fixed, and documented inadr/0016-concurrency-lock-fix.md, the same "found it, fixed it, said so" standard ADR-0011 set in Phase 5. - Integration (
tests/integration/, 9 tests, excluded from the default run): real live subprocesses — a realbilling_service, a realredis-server(this sandbox has the binary; ARQ+fakeredis is still used for the queue tier per ADR-0012), and — opt-in, credential-gated, honestly skipped without them — a real Salesforce org.
Real, measured coverage on the default (fast) tier:
86.8% statement coverage across app/ and billing_service/
(pytest --cov=app --cov=billing_service --cov-report=term-missing).
The lowest-covered files (billing_client.py, salesforce.py's live
HTTP paths, worker_settings.py) are exactly what the integration tier
and test_queue.py's real arq.Worker cover instead — no double
counting, no inflated single number claimed across tiers that test
different things.
Phase 8 — CI/CD + Deployment: done.
.github/workflows/ci-cd.yml— six gated jobs:lint-typecheck+security(parallel) →unit-tests→integration-tests→docker-build→deploy. Each stage is a realneeds:dependency; a failure anywhere stops the pipeline before it reaches deploy. Seeadr/0018-cicd-render-deployment.md.ruff check/ruff format --check/mypyall run clean — this wasn't true at the start of Phase 8 (34 lint errors, 13 type errors); see the ADR for which fixes were real bugs (anOptional[str]gap at the tool-execution boundary) versus documented, deliberate style decisions (blindexcept Exceptionat the orchestrator's fail-closed boundary).bandit(static security lint) andpip-audit(dependency CVE lookup) both run clean as a dedicated CI job — explicitly scoped as that, not claimed as full SAST/SCA coverage.- Two multi-stage, non-root Docker images
(
docker/Dockerfile,docker/Dockerfile.billing) — one per service, matching how they've run as two separate processes since Phase 1. Seeadr/0019-docker-multi-stage-build.md. render.yaml— the whole deployment topology (both services, DB, Redis, env var wiring) declared as an Infrastructure-as-Code Blueprint, not clicked together by hand. Deploys are gated on CI passing (autoDeploy: false+ a deploy-hook step that only runs after every other job succeeds), not ongit pushalone. Free tier, no card, everywhere — the documented trade-off is ~10-30s cold start after ~15 minutes idle. Seedocs/DEPLOYMENT.mdfor the full setup walkthrough andadr/0018-cicd-render-deployment.mdfor the reasoning.Makefile— every CI gate runnable locally in the same order (make ci), so a contributor can reproduce a red CI run without pushing a commit to find out why.
Real vs. Simplified (and Why)
Every simplification below was a deliberate choice, made once and named once, not discovered by a reader comparing the code to the plan. "Real" means the thing actually runs against a live dependency or enforces a real guarantee; "simplified" names the specific corner cut and why.
| Area | What's real | What's simplified | Why |
|---|---|---|---|
| Salesforce auth | A genuine OAuth2 Connected App, real token exchange, real API calls against a Developer Edition org | Username-password OAuth flow, not JWT bearer | JWT bearer needs a self-signed cert registered with the org — real added setup for a demo-scale integration with one service account. Named as a documented trade-off, not hidden. |
| Caller identity | RBAC scope + risk evaluation genuinely gates every write | caller_role is passed directly in the request body — there is no real authentication layer in front of it yet |
See adr/0008-rbac-policy-gate.md: "No real auth yet... that's a documented placeholder, not real authentication." Real auth (OAuth2 client credentials or a signed JWT per caller) is the first thing to add before this touches production traffic. |
| Idempotency | Atomic Redis SET NX EX, proven under 25 concurrent threads (adr/0009) |
Single-node Redis, no cluster/Sentinel failover story | Demo-scale; a production deployment would need Redis HA, which is infrastructure this project's $0 constraint doesn't stretch to. |
| Observability | Real Prometheus counters/histogram (GET /metrics), real structured JSON logs |
No distributed tracing (OpenTelemetry/Tempo), no log aggregation service | adr/0014-prometheus-metrics-scope.md — full tracing needs a collector + backend that isn't free-tier-friendly at this scope, and a two-service system doesn't have enough hop-to-hop complexity to make traces earn their cost yet. |
| Security scanning | bandit (static lint) + pip-audit (dependency CVEs) run in CI on every push, both currently clean |
Not a full SAST/SCA pipeline, no penetration testing, no dependency license scanning | Named explicitly in adr/0018-cicd-render-deployment.md as scoped, not comprehensive — the honest floor for a solo project, not the ceiling for what a security team would actually run. |
| PII protection | Real regex-based tokenization, applied before any context reaches the reasoning step, with a real bug found and fixed in production-like testing (an ISO timestamp misidentified as a phone number — adr/0011-pii-tokenization.md) |
Regex heuristics, not an ML-based NER model | A real trade-off: regex is auditable and has zero inference cost, but it will miss PII shapes it wasn't written for. Documented as the known failure mode, not silently accepted. |
| Deployment | Two real Docker images, a real Render Blueprint, a real CI-gated deploy hook (adr/0018, adr/0019) |
Free tier only — cold starts after ~15 min idle, no autoscaling, no multi-region | The project's own $0/no-card constraint rules out anything else; the trade-off is named plainly in docs/DEPLOYMENT.md, not hidden behind a live demo link that just happens to always be warm. |
| Type checking | mypy runs clean in CI as a real, enforced gate |
Non-strict baseline (disallow_untyped_defs = false), not --strict |
Phases 0-7 were written without type checking in the loop; flipping --strict on now would surface a large batch of pre-existing annotation gaps unrelated to any real bug. Named as incremental future work in pyproject.toml's own comments, not swept under a blanket ignore. |
| Load / throughput | Nothing claimed | No load test has been run against this project, at all | See docs/BENCHMARKS.md — stated as a gap, with what a real load-test setup would look like, rather than a fabricated req/s number. |
What was removed, not simplified — real ideas, understood, deliberately
left out of scope rather than built shallow — is its own list in
project-2-agentic-salesforce-gateway-plan.md Section 8 ("What This
Deliberately Excludes"): business-state serializability, counterfactual
simulation before execution, exception intelligence / latent-rule
discovery, an economic risk-scoring engine, full OpenTelemetry + Tempo
tracing, graph visualization of enterprise state, and any adversarial
suite beyond the three named cases in Phase 7. Each is a legitimate
"what would you build next" interview answer — none is claimed as built.
Architecture Decision Records
Every non-trivial engineering decision in this project has a written ADR —
not backfilled after the fact, one per real decision as it was made. Full
text in adr/; index below, grouped by the phase that produced it.
Foundations (Phase 0)
0001— Use FastAPI as the web framework0002— Modular monolith over microservices0003— PostgreSQL for persistent state0004— Redis for idempotency keys and task queueing0005— Use MCP for the agent's tool layer
Agent + tools (Phases 2-3)
0006— In-process MCP tool registry, not a separate MCP transport server0007— Explicit allowed-transition table for the agent state machine
Policy + security (Phases 4-5)
0008— Plain rule-based policy gate — RBAC scope, then risk level0009— Atomic Redis SET NX EX for execution idempotency0010— HMAC-SHA256 verification on approval callbacks0011— Regex-based PII tokenization before any LLM-bound context
Async + audit (Phase 6)
0012— ARQ + Redis for async case execution0013— Postgres-backed audit log, SQLite fallback for dev/test0014— Prometheus metrics — scope deliberately narrow
Testing (Phase 7)
0015— Five test tiers, not one suite0016— Per-action locking to fix a real state-machine race condition0017— Real input/output JSON schemas on every MCP tool
CI/CD + deployment (Phase 8)
0018— CI/CD pipeline shape and Render for deployment0019— Multi-stage, non-root Docker builds; two images, not one
Live Demo
Not deployed yet from this checkout. render.yaml + docs/DEPLOYMENT.md
contain everything needed to stand up a live instance on Render's free
tier in about 10 minutes once this repo is pushed to your own GitHub
account — deliberately not claiming a live URL here until one is actually
running and being kept warm, per this project's own standard of not
stating things that aren't currently true.
Once deployed, replace this section with:
Live: https://ledgent-<yours>.onrender.com/health
(Free tier — first request after ~15 min idle may take 10-30s to wake up.)
Benchmarks
Full page with every measured number and the exact command to reproduce
it: docs/BENCHMARKS.md. Headline, real, currently-passing numbers:
- 90 total test cases (81 fast + 9 integration), 8/9 integration passing, 1 honestly skipped without live Salesforce credentials.
- 86.8% statement coverage on the fast tier (
pytest --cov=...). - 25 concurrent threads, 0 errors, exactly 1 real tool call —
tests/test_concurrency.py, the tier that caught and proved the fix for the real race inadr/0016-concurrency-lock-fix.md. ruff,mypy,bandit,pip-auditall run clean in CI on every push (.github/workflows/ci-cd.yml).- No load-testing number is claimed anywhere in this repo — see
docs/BENCHMARKS.mdfor what that would take to measure honestly.
Running Locally
# 1. Copy env template and fill in your Salesforce Connected App credentials
cp .env.example .env
# 2. Start local Postgres + Redis
docker compose -f docker/docker-compose.yml up -d
# 3. Install dependencies (requirements-dev.txt pulls in requirements.txt
# plus pytest/ruff/mypy/bandit/pip-audit; use requirements.txt alone
# for a production-only install, which is what docker/Dockerfile does)
pip install -r requirements-dev.txt
# 4. Run the main app
uvicorn app.main:app --reload --port 8000
# 5. In a second terminal, run the billing service
uvicorn billing_service.main:app --reload --port 8001
# 6. In a third terminal, run the ARQ worker (Phase 6 — required for
# POST /v1/agent/cases; the /sync variant works without it)
arq app.queue.worker_settings.WorkerSettings
# 7. Run tests
pytest -v
# 8. Try the full async flow
curl -X POST localhost:8000/v1/agent/cases \
-H "Content-Type: application/json" \
-d '{"case_id":"demo-1","subscription_id":"sub_1002","caller_role":"agent"}'
# -> {"action_id": "...", "case_id": "demo-1", "status": "queued"}
curl localhost:8000/v1/agent/cases/{action_id} # poll for state
curl localhost:8000/v1/audit/{action_id} # full reconstructed timeline
curl localhost:8000/metrics # Prometheus format
Testing
# Fast tier: unit + contract + adversarial + concurrency (81 tests, ~2s)
pytest -v
# With coverage
pytest --cov=app --cov=billing_service --cov-report=term-missing
# Integration tier: spins up real live subprocesses (billing_service,
# redis-server, and app.main); Salesforce test skips honestly without
# real credentials in .env
pytest -m integration tests/integration -v
CI/CD & Deployment
Every push and PR runs the full pipeline in .github/workflows/ci-cd.yml:
lint → type check → security scan → unit tests → integration tests →
Docker build → (on main only, after everything else passes) deploy.
Run the same checks locally with make ci. See
adr/0018-cicd-render-deployment.md for why the pipeline is shaped this
way and docs/DEPLOYMENT.md for the one-time Render + GitHub setup and
live-deployment verification steps. Developing on Windows? See
docs/WINDOWS.md for PowerShell command syntax and the one real platform
gap (no native redis-server binary — handled with an honest test skip,
not a crash).
make ci # every CI gate, in CI order, stopping at the first failure
Recommended Servers
playwright-mcp
A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.
Magic Component Platform (MCP)
An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.
Audiense Insights MCP Server
Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.
VeyraX MCP
Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.
graphlit-mcp-server
The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.
Kagi MCP Server
An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.
E2B
Using MCP to run code via e2b.
Neon Database
MCP server for interacting with Neon Management API and databases
Exa Search
A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.
Qdrant Server
This repository is an example of how to create a MCP server for Qdrant, a vector search engine.