Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

2 Commits

Folders and files

Repository files navigation

Gemini Entity Extraction Latency Repro

Clean-room repro for intermittent Gemini generateContent latency outliers during entity extraction with thinking enabled.

The short version: the same entity extraction prompt usually returns in about 12-18 seconds, but intermittently returns a normal STOP response after more than 2 minutes with usageMetadata.thoughtsTokenCount around 62900. In those outlier responses, no thought text parts are returned even though thinkingConfig.includeThoughts is true.

This harness mirrors the current production request path, but the latency pattern was observed before prompt/content caching was added and has been seen on earlier Flash models as well.

See REPORT.md for the background, observed runs, and what we are trying to confirm.

Run

cat example-source.txt | GEMINI_API_KEY=... node gemini-latency-repro.mjs 20 300000 standard

Arguments:

[runs] [timeoutMs] [standard|priority|flex] [model]

Defaults:

  • runs: 1
  • timeoutMs: 60000
  • service tier: standard
  • model: gemini-3.1-flash-lite

Examples:

# Production-like timeout behavior, retries after 60s.
cat example-source.txt | GEMINI_API_KEY=... node gemini-latency-repro.mjs 20 60000 standard

# Let outliers finish so thoughtsTokenCount and response metadata are visible.
cat example-source.txt | GEMINI_API_KEY=... node gemini-latency-repro.mjs 20 300000 standard

# Compare another Gemini model slug.
cat example-source.txt | GEMINI_API_KEY=... node gemini-latency-repro.mjs 20 300000 standard gemini-3.1-flash

The script uses:

  • https://generativelanguage.googleapis.com/v1beta/cachedContents
  • https://generativelanguage.googleapis.com/v1beta/models/{model}:generateContent
  • thinkingConfig.includeThoughts: true
  • thinkingConfig.thinkingLevel: low
  • stdin source text wrapped as <source index='0'>...</source>

Output

Each successful response writes artifacts to /tmp, including:

  • gemini-latency-run-###-entity-cache-body.json
  • gemini-latency-run-###-entity-body.json
  • gemini-latency-run-###-entity-attempt-#-response.json
  • gemini-latency-run-###-entity-attempt-#-parts.json
  • gemini-latency-run-###-entity-attempt-#-thoughts.txt
  • gemini-latency-run-###-entity-attempt-#-answer.txt

The key signal to watch is a normal STOP response with very high usageMetadata.thoughtsTokenCount, high latency, and no returned thought parts.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages