# archsim — architecture simulator format

A JSON object that describes a **system architecture** (clients, gateways, services, queues, caches,
databases…) for the **Architecture Simulator** of Jarroba Tools. A person pastes it into the tool and
presses **Simulate**: traffic flows through the graph and the tool shows throughput, p50/p95 latency,
availability, queue lag, lost messages, the bottleneck and the cost per hour.

It is a *load model*, not a drawing format. Positions are computed for you; what matters is **which
components exist, how they are connected and their numbers** (capacity, latency, queue size…). If you
only need a picture of an architecture, use JTD ([jtd.md](jtd.md)) instead.

**Where it goes:** menu *Examples and templates* → **Paste JSON (AI authoring)**, or simply focus the
canvas and press **Ctrl+V**. It is **added** to the current page; nothing is deleted.

## Minimal example

```json
{ "archsim": 1, "mode": "runtime",
  "nodes": [
    { "id": "users", "kind": "client", "label": "Users", "capacity": 1000 },
    { "id": "api",   "kind": "service", "label": "Orders API", "capacity": 2000, "latency": 10 },
    { "id": "kafka", "kind": "queue", "label": "Kafka", "queue": 100000, "partitions": 6 },
    { "id": "db",    "kind": "db", "label": "Orders DB", "capacity": 400, "latency": 15 }
  ],
  "edges": [ ["users", "api"], ["api", "kafka"], ["kafka", "db"] ] }
```

Simulated, this shows the thing people get wrong about queues: the database never receives more than
it can take (400 msg/s), the other 600 msg/s pile up **in Kafka** (lag), the bar says *"full in 2 min"*,
and after that ~600 msg/s are **lost silently** — the users already got their OK.

## Top level

| field | required | meaning |
|---|---|---|
| `archsim` | yes | Format version. Always `1`. |
| `mode` | no | `"runtime"` (default). `"cicd"` exists for pipelines but is not described here. |
| `nodes` | yes | Non-empty list of components (below). |
| `edges` | no | Connections, each `["fromId", "toId"]` or `{ "from": "…", "to": "…" }`. Direction = direction of the requests. |

## Node fields

Every field except `id` and `kind` is optional; **leave out what you do not know** and the type's
default is used (table further down).

| field | type | meaning |
|---|---|---|
| `id` | string | Unique within the document. Used by `edges`. |
| `kind` | string | Component type (table below). An unknown kind is an error. |
| `label` | string | Name shown on the canvas. |
| `capacity` | number | Requests per second it can serve. For `client`/`attacker`: requests per second it **emits**. For `ratelimit`: the allowed rate. |
| `latency` | number | Mean service time, ms. |
| `jitter` | number | ± variation of the service time, ms. Only inflates p95, and only near saturation. |
| `queue` | number | Max. requests waiting before dropping. `0` = unbounded. For `queue`/`pubsub` it is the **backlog cap**. |
| `fanout` | `"split"` \| `"all"` | How traffic leaves the node. `split`: each request goes to ONE output (load balancing; outputs that are down are skipped). `all`: the node **calls every** output. Default: `split` for `client`, `attacker`, `lb`, `gateway`, `waf`, `ratelimit`, `cdn`, `queue`; `all` for everything else. |
| `hitRatio` | 0–1 | `cache`/`cdn` only. Fraction of hits; only misses continue downstream (default 0.8 / 0.9). |
| `burst` | number | `ratelimit` only. Token-bucket size (burst admitted). Default = one second of rate. |
| `replicas` | integer | `db`/`nosql`/`search` only. Read replicas. |
| `readRatio` | 0–1 | With `replicas`: fraction of reads (default 0.8). Writes all go to the primary. |
| `coldStartMs` | number | `service`/`worker`/`auth`. > 0 makes it **serverless**: scales by itself, but capacity beyond the warm instances pays this cold start. |
| `partitions` | integer | `queue`/`pubsub` only (Kafka). Each partition is consumed serially by ONE consumer, so the group cannot exceed `partitions × 1000 / consumer latency`. `0` = no ceiling (SQS, RabbitMQ). |
| `redelivery` | 0–100 | `queue`/`pubsub` only. % of messages delivered more than once (at-least-once). |
| `idempotent` | boolean | On a **consumer**: it deduplicates, so redeliveries do not amplify its load. |

## Kinds

Defaults are for a node that sets nothing. Capacity in req/s, latency in ms.

| kind | what it is | capacity | latency | queue |
|---|---|---|---|---|
| `client` | Source of legitimate traffic | 100 (emits) | 0 | – |
| `attacker` | Source of attack traffic (DDoS) | 5000 (emits) | 0 | – |
| `waf` | WAF / filter: blocks ~90 % of attack traffic | 20000 | 2 | 500 |
| `gateway` | API gateway (routes: `split`) | 2000 | 2 | 200 |
| `lb` | Load balancer (`split`, skips backends that are down) | 5000 | 1 | 500 |
| `cdn` | CDN (hit ratio 0.9) | 20000 | 5 | 1000 |
| `ratelimit` | Rate limiter, token bucket; rejects at once (429) | 8000 | 1 | – |
| `service` | Service / microservice | 200 | 20 | 100 |
| `worker` | Background worker | 100 | 50 | 100 |
| `auth` | Auth / identity | 1000 | 15 | 100 |
| `orchestrator` | Orchestrator (saga) | 500 | 10 | 200 |
| `queue` | Queue (Kafka, SQS…): asynchronous, consumers compete | 10000 (ingest) | 1 | 100000 |
| `pubsub` | Event bus: every subscriber gets every event | 15000 (ingest) | 1 | 100000 |
| `cache` | Cache (Redis), hit ratio 0.8 | 8000 | 1 | 200 |
| `db` | SQL database | 400 | 15 | 100 |
| `nosql` | NoSQL | 2000 | 8 | 200 |
| `objstore` | Object store (S3) | 3000 | 30 | 500 |
| `search` | Search (Elastic) | 1500 | 20 | 150 |
| `etl` | ETL / ingestion | 200 | 500 | 1000 |
| `datalake` | Data lake | 5000 | 40 | 500 |
| `lakehouse` | Lakehouse | 800 | 120 | 200 |
| `warehouse` | Data warehouse | 300 | 200 | 100 |
| `bi` | BI / dashboard | 100 | 150 | 50 |
| `external` | Third-party API | 100 | 120 | 50 |
| `agent` | AI agent (generic) | 50 | 100 | 50 |
| `llm` | LLM (generic, self-hosted) | 20 | 800 | 40 |
| `tool` | Tool / MCP server | 100 | 40 | 50 |
| `vectordb` | Vector DB | 500 | 25 | 100 |
| `embed` | Embeddings | 200 | 60 | 60 |

The tool also has detailed AI-agent kinds (`ai_*`) and CI/CD kinds (`ci_*`); they import with their
defaults only, because their specific settings are not part of this format.

## What the simulation does (so your numbers mean something)

- **A service calls all its dependencies; a balancer picks one.** `service → auth, cache, db` means
  every request uses the three. `lb → a, b` means half goes to each.
- **Queues are asynchronous.** The producer gets its OK when the message is accepted; consumers pull
  at their own pace. On average a queue does **not** remove load, it defers it: sustained excess
  grows the backlog until `queue`, then messages are lost silently. Queues are for **peaks** and for
  surviving a consumer outage. The response latency ends at the queue; the processing latency
  includes the lag.
- **Caches cut the path.** Read-through (`cache → db`) and cache-aside (`service → cache` **and**
  `service → db`) both send only the misses to the database. p50/p95 follow the paths requests
  actually take: with 80 % hits, p50 does not include the database; with 1 % misses, not even p95.
- **Rate limiters reject, they do not queue.** Excess over the rate is a 429, not a wait.
- **Capacity is per node.** A database with capacity 400 behind a service with 2000 is the
  bottleneck no matter how the service is tuned.

## Mistakes that matter

1. **Drawing the response path backwards.** Edges follow the **requests** (client → … → db), not the
   data coming back.
2. **Putting a queue in front of something to "reduce its load".** It only moves the excess to the
   backlog. If average input > consumer capacity, add consumers, partitions or capacity.
3. **Linking a service to several databases expecting them to share the load.** A service calls
   ALL its outputs. For shards, set `"fanout": "split"` on the service or put an `lb` in between.
4. **Inventing precise numbers.** If you do not know a capacity, leave it out: the default is a
   reasonable small instance. A made-up number looks authoritative and nobody checks it.
5. **Returning prose or code fences around the JSON.** It is pasted verbatim.

## How to check your work

You cannot judge a load model by reading the JSON. Paste it, press **Simulate** (16× makes queues
fill in seconds) and look at the status bar: throughput, p50/p95, availability, and — if there are
queues — lag, *"full in…"* and lost messages. Then open **Report**: it lists the bottleneck, queues
that grow, databases without a cache, retries without a circuit breaker, and the cost. Fix and paste
again.

**Over MCP** (local, nothing leaves the machine): `archsim_validate` checks the document, and
`archsim_simulate` runs it with the same engine and returns steady-state throughput, p50/p95,
availability, bottleneck, cost, per-queue lag / *"full in"* / lost messages, and the Report's
findings. Pass `phase: {"seconds": 60, "scenario": {"loadMult": 6}}` to see a peak — autoscaling takes
about a minute to arrive, so a short peak is exactly where designs fail.
