Clean-room repro for intermittent Gemini generateContent latency outliers during entity extraction with thinking enabled.
The short version: the same entity extraction prompt usually returns in about 12-18 seconds, but intermittently returns a normal STOP response after more than 2 minutes with usageMetadata.thoughtsTokenCount around 62900. In those outlier responses, no thought text parts are returned even though thinkingConfig.includeThoughts is true.
This harness mirrors the current production request path, but the latency pattern was observed before prompt/content caching was added and has been seen on earlier Flash models as well.
See REPORT.md for the background, observed runs, and what we are trying to confirm.
cat example-source.txt | GEMINI_API_KEY=... node gemini-latency-repro.mjs 20 300000 standardArguments:
[runs] [timeoutMs] [standard|priority|flex] [model]
Defaults:
runs:1timeoutMs:60000- service tier:
standard - model:
gemini-3.1-flash-lite
Examples:
# Production-like timeout behavior, retries after 60s.
cat example-source.txt | GEMINI_API_KEY=... node gemini-latency-repro.mjs 20 60000 standard
# Let outliers finish so thoughtsTokenCount and response metadata are visible.
cat example-source.txt | GEMINI_API_KEY=... node gemini-latency-repro.mjs 20 300000 standard
# Compare another Gemini model slug.
cat example-source.txt | GEMINI_API_KEY=... node gemini-latency-repro.mjs 20 300000 standard gemini-3.1-flashThe script uses:
https://generativelanguage.googleapis.com/v1beta/cachedContentshttps://generativelanguage.googleapis.com/v1beta/models/{model}:generateContentthinkingConfig.includeThoughts: truethinkingConfig.thinkingLevel: low- stdin source text wrapped as
<source index='0'>...</source>
Each successful response writes artifacts to /tmp, including:
gemini-latency-run-###-entity-cache-body.jsongemini-latency-run-###-entity-body.jsongemini-latency-run-###-entity-attempt-#-response.jsongemini-latency-run-###-entity-attempt-#-parts.jsongemini-latency-run-###-entity-attempt-#-thoughts.txtgemini-latency-run-###-entity-attempt-#-answer.txt
The key signal to watch is a normal STOP response with very high usageMetadata.thoughtsTokenCount, high latency, and no returned thought parts.