Concurrency and Throughput
Introduction
A single drawing takes as long as it takes. Reading a simple sheet is a matter of seconds; a dense drawing with many features can run close to a minute, and no amount of tuning on your side changes that number.
Throughput is a different quantity, and it is one you control. Requests are independent of one another, so the way to read a thousand drawings faster is to have more than one of them in flight at a time. This page describes what happens when you do.
One Request, One Drawing
Each request carries exactly one drawing and any number of Asks. Asks within a request share the expensive preparation work, so asking for META_DATA and FEATURES in one request is considerably cheaper and faster than asking for them in two.
Requests, by contrast, share nothing. Submitting two of them concurrently is the same as two separate clients submitting one each, and that is what makes concurrency the lever for throughput.
Group your Asks, spread your requests
Put everything you want to know about one drawing into a single request. Put different drawings into different requests, and run those requests concurrently.
The First Request After an Idle Period
The API layer scales on demand. When it has been idle, the first request has to bring capacity up before it can do anything else, which adds roughly one to two and a half seconds to that request and to nothing after it.
Werk24 runs a scheduled probe that keeps a warm instance available, so a client that submits at a steady rate will not normally see this at all. Two situations still can:
- A burst launched from cold. Concurrent requests cannot share one warm instance, so if you go from idle to twenty requests at once, the ones beyond the warm instance pay the start-up cost. It is paid once per instance, not once per request, so the effect fades within the first few seconds of the batch.
- A very sparse workload. One request every few hours behaves like a burst of one, every time.
If a consistent time-to-first-result matters more to you than it costs, submit a small warm-up request before the batch, or keep your submission rate steady rather than bursty. For anything running overnight, the start-up cost is a rounding error on the total and is not worth engineering around.
Priority and Queueing
Concurrency decides how many of your requests reach the platform at once. Priority decides the order in which the platform works through what has reached it.
The two combine in a way worth planning for: a large batch submitted at your account's default priority competes with your own interactive requests. Submit batches at PRIO3 and they queue behind your real-time work instead.
See Processing Priority for the full picture, including the priority levels your account allows.
Choosing a Concurrency Level
Your account has a number of slots: how many of your requests are worked on at the same time. Best-effort and Priority include one slot and Express three; more, up to ten, are an add-on on Production (and on the retired Essential and Growth plans), or part of a contract, listed on werk24.io/pricing. The slots follow the priority of each request: a request sent at a lower priority than your account's gets that priority's slots, unless you booked more. A request beyond your slots is not refused. It waits behind your own earlier requests, and other customers' requests are not held up by it.
There is no single right number, but the shape of the trade-off is stable:
| Concurrency | Behaviour |
|---|---|
| 1 | Total time is the sum of every drawing. Simple, and slow for a batch. |
| Up to your slots | Requests run side by side. With several slots, this is where a batch gets faster. |
| Beyond your slots | No faster: the extra requests wait for one of your slots. With one slot, 8 in flight finish no sooner than 1 at a time. |
Start at the low end, measure, and raise it only while the total time is still falling. Unbounded concurrency (submitting every drawing in a directory at once) is the common mistake: it does not finish sooner, and it turns a clean batch into an exercise in retry handling.
Bounding Concurrency in Python
asyncio.Semaphore is the least intrusive way to put a ceiling on it:
Two details in there matter:
- One client, reused. The client holds the connection to the API. Creating one per drawing pays the connect and the authorization on every drawing.
return_exceptions=True. Without it, the first drawing that raises ends thegatherimmediately and that exception reaches you, so the results of every other drawing are lost even though those reads keep running in the background. With it,gatherwaits for all of them and hands back a list in which a failed drawing is represented by its exception object, so one bad file costs you that file and nothing else:
Large Batches: Use Callbacks
Holding hundreds of concurrent connections open for the duration of a batch is a poor use of both ends. For anything at that scale, submit the drawings with read_drawing_with_callback and let the results arrive at an endpoint of yours as they are produced. The submission then costs a single short call per drawing, nothing has to stay connected, and a client that restarts mid-batch loses nothing.
A callback request carries the drawing inside its body, so it takes drawings up to about 4.6 MB (see Drawing File Size Limit). If a batch can hold larger drawings, send those with read_drawing. From versions newer than 2.7.0 the client raises CallbackDrawingTooLargeException for them before anything is sent, so you can catch it and route that drawing to read_drawing instead.
Keep the request_id: it is what joins the callback you receive back to the drawing you sent. See Reference ID for attaching your own identifier instead.
A 429 Is Not Backpressure
The API does not answer a wide batch with 429. The only 429 it sends is QUOTA_EXHAUSTED: the account's request allowance is used up. That does not reset with time and names no time to wait, so waiting and retrying sends the same refusal. Stop the batch and contact your Werk24 account team. With read_drawing the refusal does not arrive as a 429: the read fails at its start (see HTTP Error Codes).
See HTTP Error Codes for the full response shape and the other statuses worth handling in a batch.
See Also
- Processing Priority - queueing order and the priority override
- Hooks - handling results as they stream in, and keeping the handler small enough not to become the bottleneck
- Python Client - synchronous, asynchronous and callback-based implementations
- Webhooks - registering a callback endpoint
- HTTP Error Codes - complete error handling reference