Skip to content

AskDocumentProfile

Description

AskDocumentProfile returns fast metadata about the document you sent and how long it is likely to take: the number of pages, the sheet format, and a processing-time estimate.

It is answered from the file's shape before any interpretation begins and without asking a model anything, so it arrives almost immediately and costs nothing extra. Ask for it next to the asks you actually need, and it is the first result you receive.

When to use

  • Show a waiting user how long they are waiting. Reading a drawing takes seconds to minutes. With the profile, a progress bar or an ETA can replace an open-ended spinner.
  • Size timeouts. Use processing_time.seconds_p95 rather than a fixed number.
  • Decide early whether to keep waiting, for example on a document with many more pages than you expected.

Show seconds_p50, size against seconds_p95

The estimate is a range and not a single number, because processing time has a long right tail: the median is around half a minute while the 99th percentile is over two minutes. Show seconds_p50 to a user; size timeouts and progress bars against seconds_p95.

processing_time is null when the document's shape is too unusual to compare against anything measured. Treat that as "no estimate", not as "fast".

A few things the fields do not say:

  • page_count is the length of the whole document, before any page limit is applied. It can therefore be larger than the number of pages that are read.
  • paper_size is null for inputs that state no physical size, which is every raster image: a photograph or a scan carries pixels, not millimetres. CUSTOM is a real sheet that is not one of the named formats.
  • What kind of page the document holds is not part of this ask. Request AskPageAssessment for the page type and for whether the pages carry welding callouts. The two are separate asks on purpose: asking for both gets you the profile straight away and the assessment later, rather than making the fast answer wait for the slow one.

Example Usage

1
2
3
4
5
6
from werk24 import AskDocumentProfile, AskType, read_example_drawing

results = read_example_drawing([AskDocumentProfile()])

profile = results[AskType.DOCUMENT_PROFILE][0].payload_dict
print(profile.page_count, profile.paper_size, profile.processing_time)

To use the estimate while the read is still running, iterate over the messages as they arrive. The profile comes first:

import asyncio

from werk24 import (
    AskDocumentProfile,
    AskFeatures,
    AskType,
    TechreadMessageType,
    Werk24Client,
)


async def read_with_eta(path: str) -> None:
    asks = [AskDocumentProfile(), AskFeatures()]
    async with Werk24Client() as client:
        with open(path, "rb") as drawing:
            async for message in client.read_drawing(drawing, asks):
                if message.message_type != TechreadMessageType.ASK:
                    continue
                if message.message_subtype == AskType.DOCUMENT_PROFILE:
                    estimate = message.payload_dict.processing_time
                    if estimate is not None:
                        print(f"Usually ready in about {estimate.seconds_p50:.0f} s")
                elif message.message_subtype == AskType.FEATURES:
                    print("Features:", message.payload_dict)


asyncio.run(read_with_eta("drawing.pdf"))

ResponseDocumentProfile

Bases: Response

The document's shape: how big it is, and how long it is likely to take.

Read off the file itself, before any interpretation begins and without asking a model anything, so it arrives in milliseconds rather than seconds. Use it to tell a waiting user how long they are waiting.

What kind of page this is is no longer here. It costs a vision call, which is thousands of times slower than everything in this response, so keeping the two together meant the fast answer waited for the slow one. Ask AskPageAssessment for it, and for whether the pages carry welding.

PARAMETER DESCRIPTION
ask_version

TYPE: Literal['v2'] DEFAULT: 'v2'

ask_type

TYPE: Literal[<AskType.DOCUMENT_PROFILE: 'DOCUMENT_PROFILE'>] DEFAULT: <AskType.DOCUMENT_PROFILE: 'DOCUMENT_PROFILE'>

page_count

Number of pages in the document, before any page limit is applied. Note that this does NOT drive the processing-time estimate: measured over 4,639 requests, two-page documents come back faster than one-page ones, so the estimate is keyed on sheet size instead.

TYPE: int

paper_size

The sheet format. CUSTOM for a real sheet that is not one of the named formats. None for inputs that state no physical size, which is every raster image: a photograph or a scan carries pixels, not millimetres.

TYPE: PaperSize | None DEFAULT: None

processing_time

How long this document is likely to take. None when the shape is too unusual to compare against anything we have measured; treat that as 'no estimate' rather than 'fast'.

TYPE: ProcessingTimeEstimate | None DEFAULT: None

Source code in werk24/models/v2/responses.py
class ResponseDocumentProfile(Response):
    """
    The document's shape: how big it is, and how long it is likely to take.

    Read off the file itself, before any interpretation begins and without
    asking a model anything, so it arrives in milliseconds rather than
    seconds. Use it to tell a waiting user how long they are waiting.

    **What kind of page this is is no longer here.** It costs a vision call,
    which is thousands of times slower than everything in this response, so
    keeping the two together meant the fast answer waited for the slow one.
    Ask `AskPageAssessment` for it, and for whether the pages carry welding.
    """

    ask_type: Literal[AskType.DOCUMENT_PROFILE] = AskType.DOCUMENT_PROFILE

    page_count: int = Field(
        ...,
        ge=1,
        description=(
            "Number of pages in the document, before any page limit is "
            "applied. Note that this does NOT drive the processing-time "
            "estimate: measured over 4,639 requests, two-page documents come "
            "back faster than one-page ones, so the estimate is keyed on "
            "sheet size instead."
        ),
        examples=[1, 12],
    )
    paper_size: Optional[PaperSize] = Field(
        None,
        description=(
            "The sheet format. CUSTOM for a real sheet that is not one of the "
            "named formats. None for inputs that state no physical size, "
            "which is every raster image: a photograph or a scan carries "
            "pixels, not millimetres."
        ),
        examples=[PaperSize.A3, PaperSize.ANSI_D, PaperSize.CUSTOM],
    )
    processing_time: Optional[ProcessingTimeEstimate] = Field(
        None,
        description=(
            "How long this document is likely to take. None when the shape is "
            "too unusual to compare against anything we have measured; treat "
            "that as 'no estimate' rather than 'fast'."
        ),
    )

ProcessingTimeEstimate

Bases: BaseModel

How long this document is likely to take to read.

Deliberately a range and not a single number. Werk24's processing time has a long right tail — the median is around half a minute while the 99th percentile is over two — so a point estimate would be wrong in the only case where being wrong is expensive: the request you are still waiting on. Show seconds_p50 to a user; size timeouts and progress bars against seconds_p95.

The estimate is made from the document's sheet size at the moment the file is read, before any interpretation, so it is available almost immediately and does not depend on what the drawing turns out to contain.

Deliberately not from page count, which is the obvious candidate and is wrong: measured over 4,639 requests, two-page documents come back faster than one-page ones (median 9.3 s against 18.7 s), almost certainly because a two-page PDF is usually a drawing plus a cover. Scaling by it would make the estimate worse.

PARAMETER DESCRIPTION
seconds_p50

Median expected processing time in seconds. Half of comparable documents finish faster than this.

TYPE: float

seconds_p95

95th-percentile expected processing time in seconds. Size timeouts against this rather than against the median.

TYPE: float

Source code in werk24/models/v2/models.py
class ProcessingTimeEstimate(BaseModel):
    """How long this document is likely to take to read.

    Deliberately a **range and not a single number**. Werk24's processing time
    has a long right tail — the median is around half a minute while the 99th
    percentile is over two — so a point estimate would be wrong in the only
    case where being wrong is expensive: the request you are still waiting on.
    Show `seconds_p50` to a user; size timeouts and progress bars against
    `seconds_p95`.

    The estimate is made from the document's **sheet size** at the moment the
    file is read, before any interpretation, so it is available almost
    immediately and does not depend on what the drawing turns out to contain.

    Deliberately not from page count, which is the obvious candidate and is
    wrong: measured over 4,639 requests, two-page documents come back faster
    than one-page ones (median 9.3 s against 18.7 s), almost certainly because
    a two-page PDF is usually a drawing plus a cover. Scaling by it would make
    the estimate worse.
    """

    seconds_p50: float = Field(
        ...,
        gt=0,
        description=(
            "Median expected processing time in seconds. Half of comparable "
            "documents finish faster than this."
        ),
        examples=[16.5, 34.9],
    )
    seconds_p95: float = Field(
        ...,
        gt=0,
        description=(
            "95th-percentile expected processing time in seconds. Size "
            "timeouts against this rather than against the median."
        ),
        examples=[48.0, 120.0],
    )

    @model_validator(mode="after")
    def _percentiles_are_ordered(self) -> "ProcessingTimeEstimate":
        """A 95th percentile below the median is not a distribution.

        Checked because this object exists to be acted on: a caller sizing a
        timeout against `seconds_p95` would silently get a shorter one than
        the median it is meant to bound, and nothing downstream would notice.

        Equality is allowed. A sheet size with few observations can genuinely
        report the same value at both percentiles, and rejecting that would
        turn a thin distribution into an error.
        """
        if self.seconds_p95 < self.seconds_p50:
            raise ValueError(
                f"seconds_p95 ({self.seconds_p95}) is below seconds_p50 "
                f"({self.seconds_p50}); percentiles must not decrease"
            )
        return self

Example Response

One response for the whole document:

{
    "ask_version": "v2",
    "ask_type": "DOCUMENT_PROFILE",
    "page_count": 1,
    "paper_size": "A3",
    "processing_time": {
        "seconds_p50": 16.5,
        "seconds_p95": 48.0
    }
}