Curious how the switch felt day-to-day for a small team — was the win mostly lower spend, or did the outputs stay predictable enough that you stopped second-guessing every call?
We have Jev running as shadow calls alongside with the control calls (LLM models). Agreement rate differ across use cases. We then use this data to tune the confidence and probabilities threshold for each case to achieve around high enough agreement rate with the current model.
Curious how the switch felt day-to-day for a small team — was the win mostly lower spend, or did the outputs stay predictable enough that you stopped second-guessing every call?
We have Jev running as shadow calls alongside with the control calls (LLM models). Agreement rate differ across use cases. We then use this data to tune the confidence and probabilities threshold for each case to achieve around high enough agreement rate with the current model.
I’m the author, tysm for sharing the article! Happy to share more if anyone has questions.
so much jev hype, honestly deserved