<!-- library-gap-audit: sglang-offline-engine -->
Summary
sglang is a high-performance LLM/multimodal serving framework with RadixAttention, continuous batching, and speculative decoding, positioned alongside vllm (already tracked as a separate open instrumentation gap in this repo) as one of the two dominant open-source inference engines. This repository has zero instrumentation for it.
- Current version 0.5.20 (released 2026-09-18), under very active development
- Not present in
py/pyproject.toml [tool.braintrust.matrix], no noxfile.py session, no directory under py/src/braintrust/integrations/ or py/src/braintrust/wrappers/
- Repo-wide grep for
sglang under py/ returns zero matches
What is missing
SGLang's in-process Offline Engine API (sglang.Engine, docs: https://docs.sglang.io/docs/basic_usage/offline_engine_api) is the primary way Python applications call it directly (as opposed to going through its OpenAI-compatible HTTP server, which a user could separately point the already-instrumented openai client at). None of these produce a Braintrust span today:
Engine.generate() — synchronous text/chat generation, non-streaming and streaming variants
Engine.async_generate() — asyncio generation entrypoint, non-streaming and streaming variants; this is the method the framework's own example server-wrapping code (examples/runtime/engine/custom_server.py) uses to build custom serving layers
Engine.encode() — embedding execution (e.g. BERT-base/e5-mistral style models) for retrieval/similarity use cases
Engine.rerank() — query/document relevance scoring, analogous to the cohere.rerank span type this repo already instruments (py/src/braintrust/integrations/cohere/patchers.py)
This engine also supports vision-language model (VLM) inference via the same generate()/async_generate() surface with image inputs.
Weekly downloads
Weekly downloads: 574,007 (as of 2026-09-21; https://pypistats.org/api/packages/sglang/recent)
Braintrust docs status
not_found — checked https://www.braintrust.dev/docs/integrations/ai-providers and https://www.braintrust.dev/docs/guides/traces/integrations; neither mentions SGLang. (vllm, the closest analog, is also absent from both pages and is tracked separately.)
Upstream sources
Local repo files inspected
py/src/braintrust/integrations/ — no sglang/ directory
py/src/braintrust/wrappers/ — no SGLang wrapper
py/pyproject.toml [tool.braintrust.matrix] — no sglang entry
py/noxfile.py — no test_sglang session
- Repo-wide case-insensitive grep for
sglang under py/ — zero matches
<!-- library-gap-audit: sglang-offline-engine -->
Summary
sglangis a high-performance LLM/multimodal serving framework with RadixAttention, continuous batching, and speculative decoding, positioned alongsidevllm(already tracked as a separate open instrumentation gap in this repo) as one of the two dominant open-source inference engines. This repository has zero instrumentation for it.py/pyproject.toml[tool.braintrust.matrix], nonoxfile.pysession, no directory underpy/src/braintrust/integrations/orpy/src/braintrust/wrappers/sglangunderpy/returns zero matchesWhat is missing
SGLang's in-process Offline Engine API (
sglang.Engine, docs: https://docs.sglang.io/docs/basic_usage/offline_engine_api) is the primary way Python applications call it directly (as opposed to going through its OpenAI-compatible HTTP server, which a user could separately point the already-instrumentedopenaiclient at). None of these produce a Braintrust span today:Engine.generate()— synchronous text/chat generation, non-streaming and streaming variantsEngine.async_generate()— asyncio generation entrypoint, non-streaming and streaming variants; this is the method the framework's own example server-wrapping code (examples/runtime/engine/custom_server.py) uses to build custom serving layersEngine.encode()— embedding execution (e.g. BERT-base/e5-mistral style models) for retrieval/similarity use casesEngine.rerank()— query/document relevance scoring, analogous to thecohere.rerankspan type this repo already instruments (py/src/braintrust/integrations/cohere/patchers.py)This engine also supports vision-language model (VLM) inference via the same
generate()/async_generate()surface with image inputs.Weekly downloads
Weekly downloads: 574,007 (as of 2026-09-21; https://pypistats.org/api/packages/sglang/recent)
Braintrust docs status
not_found— checked https://www.braintrust.dev/docs/integrations/ai-providers and https://www.braintrust.dev/docs/guides/traces/integrations; neither mentions SGLang. (vllm, the closest analog, is also absent from both pages and is tracked separately.)Upstream sources
Local repo files inspected
py/src/braintrust/integrations/— nosglang/directorypy/src/braintrust/wrappers/— no SGLang wrapperpy/pyproject.toml[tool.braintrust.matrix]— nosglangentrypy/noxfile.py— notest_sglangsessionsglangunderpy/— zero matches