Skip to content
Discussion options

You must be logged in to vote

Hey! Just released a new version which includes performance improvements that should help with this.

Here's what's new:

  • Thinking mode toggles — You can now disable thinking/reasoning for your LLM in the settings, which can significantly speed up responses.
  • Faster Whisper transcription — CPU transcription has been optimised, so voice input processing is quicker.

Beyond that, response speed mainly depends on which LLM model you're running and your hardware. A few tips:

  1. Use a smaller model — If you're running a large model (e.g. 14B+ parameters), try a smaller one. Something in the 3B-7B range will respond much faster.
  2. Disable thinking mode — In settings, turn off thinking/extended thinki…

Replies: 1 comment

Comment options

You must be logged in to vote
0 replies
Answer selected by isair
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment
Category
Q&A
Labels
None yet
2 participants