What the 87ms figure means
The 87ms figure is P50 end-to-end request latency for Forge on owned NVIDIA DGX hardware with no shared queue. It measures the complete request path, including the network round trip. It is not time to first byte.
Measurement boundary
The measurement begins when the client sends a request and ends when the client receives the completed response. Reporting the full request path keeps the number aligned with what an application or agent experiences.
The 300ms figure is an observed market reference across hosted embedding APIs. It is context, not an independent benchmark or a claim that every hosted API has the same latency.
Reproducing the request path
Use the public Playground to exercise the hosted request path without an account. Production latency varies with network location, payload size, selected tier, and deployment topology.
The infrastructure and model provenance behind Forge are documented in Ingot Poured.