@wochinge

Product Eng @ Langfuse | Coding, bouldering, coffee, Bayern Munich | he/him

Munich, Germany
Joined July 2009
LLM-as-a-judge evaluators in Langfuse just became a lot more powerful! Multi-modal support: They can now actually access multi-modal content in your observations to check whether the agent behaved correctly
llm-as-a-judge evaluators now support multi-modal inputs. score the images, audio, video, PDFs, and text files captured in your observations. prompts can be multi-message too: system for the criteria, user for the content. langfuse.com/changelog/2026-…
1
1
51
Make your evaluator more robust by leverage multi-message prompts: - `system`: for the judge instructions - `user`: for the evaluated content - `assistant`: for any examples that you give the judge
8
Tobias Wochinger retweeted
llm-as-a-judge evaluators now support multi-modal inputs. score the images, audio, video, PDFs, and text files captured in your observations. prompts can be multi-message too: system for the criteria, user for the content. langfuse.com/changelog/2026-…
11
10
3
56
8,372
God - shipping from my phone is addicting. Expect more changelogs for evals in Langfuse in the next days!
1
22
We completely rebuilt how evaluators are set up and managed! How we did it? Shipping a 50k line PR on a Saturday! Why? Check the thread!
Evals Evals Evals 𝗦𝗲𝘁𝘁𝗶𝗻𝗴 𝘂𝗽 𝗲𝘃𝗮𝗹𝘂𝗮𝘁𝗼𝗿𝘀 𝗶𝗻 𝗟𝗮𝗻𝗴𝗳𝘂𝘀𝗲 𝗵𝗮𝘀 𝗻𝗲𝘃𝗲𝗿 𝗯𝗲𝗲𝗻 𝗲𝗮𝘀𝗶𝗲𝗿 Choose from templates to get started, select real production data to test your evaluators on and use the interactive variable mapping to get the right data points into the evaluation context. Watch @wochinge's walk-through and check out the new UX
1
2
67
and that we released during a time when we knew interactions with evaluators would be minimal.
1
3
As exciting as the release was: I'm even more excited what we can build on this new foundation! Things like multi-modal support or cost alerts are already on the underway! Try out the new eval UX and let me know what you think!
3
Code evaluators were my first big feature at Langfuse. Adoption is growing every day, and it has been really rewarding to see how many people already use. And also: How many immediately tried to break the tenant isolation.
1
4
7
574
It was clear that as soon as we'd launch it people would be all over us. Great to work in an environment that trusts new engineers to work on sth so crucial within their first weeks - but definitely nothing that I wanted to get wrong.
1
24
In above blog post I share in full depth how we went about the implementation, how the isolation model works and how the first real life encounters went.
19
Did this on Langfuse’s CI recently. Told Fable to make our pipelines 30s faster end-to-end, iterating on real CI timings with old runs as the baseline. 10% speed up over night - quite the win for an already very fast pipeline (5min).
I'm having a lot of success giving GPT more ambitious /goal criteria, even if I don't expect the model to ever hit the goal. E.g., "Improve performance by 10%" rather than 1%. Tell it to be ambitious, think big, etc. Then put up PRs as it goes and merge discrete improvements.
1
1
87
Our CI is already fast (~5min) so wins are non-obvious. I Optimizing CI is normally tedious with having 2 wait between attempts. Fable just pushed through: thesis, test run, repeat until it reached the goal.
1
1
44
Comparing runs a week later, the improvements nicely held up with about 33s savings across runs 🎉
9