fae53369bf
Switch default Gemini model to 3.1 Flash-Lite preview ( #6625 )
...
3.1 Flash-Lite is faster (throughput + TTFT), meaningfully smarter, and
still among the cheapest Gemini options, making it a strict upgrade
over 2.5 Flash-Lite for narrative text generation.
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-04-09 19:00:51 -07:00
89ad8371b5
Add Gemini LLM support with model comparison documentation ( #5561 )
...
This adds Google Gemini as a third LLM provider option alongside OpenAI
and Anthropic (Claude). Based on streaming latency tests, Gemini 2.5
Flash-Lite shows the fastest time-to-first-token (~0.6s) at the lowest
cost ($0.10/$0.40 per 1M tokens).
Changes:
- Add GeminiServiceImpl.scala following OpenAI/Claude pattern
- Add GeminiModelName setting (default: gemini-2.5-flash-lite)
- Update LlmResolver to support "gemini" provider
- Update admin console dropdowns with Gemini options
- Reorder dropdown options to show recommended models first:
- OpenAI: gpt-4.1-mini (was gpt-5-mini)
- Claude: claude-3-5-haiku (already first)
- Gemini: gemini-2.5-flash-lite
- Add GEMINI_API_KEY to docker-compose and deployment workflow
- Add docs/LLM_MODEL_COMPARISON.md with latency/pricing comparison
Streaming TTFT results from testing:
- Gemini 2.5 Flash-Lite: ~0.6s (fastest, cheapest)
- gpt-4.1-mini: ~1.7s
- claude-3-5-haiku: ~1.9s
- Gemini 2.5 Flash: ~6.8s (surprisingly slow)
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com >
2026-01-23 08:23:38 -08:00