@wochingei
iAccount based inGermany
About this account
- Account based in
- Germany
- Connected via
- Germany App Store
Account-level information from X, not a live location or the device used for a specific post.
Product Eng @ Langfuse | Coding, bouldering, coffee, Bayern Munich | he/him
Munich, Germany
Joined July 2009
- Tweets171
- Following347
- Followers70
- Likes2.3K
LLM-as-a-judge evaluators in Langfuse just became a lot more powerful!
Multi-modal support: They can now actually access multi-modal content in your observations to check whether the agent behaved correctly
llm-as-a-judge evaluators now support multi-modal inputs. score the images, audio, video, PDFs, and text files captured in your observations.
prompts can be multi-message too: system for the criteria, user for the content.
langfuse.com/changelog/2026-…
Tobias Wochinger retweeted
llm-as-a-judge evaluators now support multi-modal inputs. score the images, audio, video, PDFs, and text files captured in your observations.
prompts can be multi-message too: system for the criteria, user for the content.
langfuse.com/changelog/2026-…
God - shipping from my phone is addicting. Expect more changelogs for evals in Langfuse in the next days!
We completely rebuilt how evaluators are set up and managed!
How we did it? Shipping a 50k line PR on a Saturday!
Why? Check the thread!
Evals Evals Evals
𝗦𝗲𝘁𝘁𝗶𝗻𝗴 𝘂𝗽 𝗲𝘃𝗮𝗹𝘂𝗮𝘁𝗼𝗿𝘀 𝗶𝗻 𝗟𝗮𝗻𝗴𝗳𝘂𝘀𝗲 𝗵𝗮𝘀 𝗻𝗲𝘃𝗲𝗿 𝗯𝗲𝗲𝗻 𝗲𝗮𝘀𝗶𝗲𝗿
Choose from templates to get started, select real production data to test your evaluators on and use the interactive variable mapping to get the right data points into the evaluation context.
Watch @wochinge's walk-through and check out the new UX
Code evaluators were my first big feature at Langfuse.
Adoption is growing every day, and it has been really rewarding to see how many people already use.
And also: How many immediately tried to break the tenant isolation.
code evaluators means running untrusted user code across thousands of tenants in our multi-tenant cloud.
@wochinge wrote about the full design: (V8 isolates, pyodide), gVisor and firecracker, and kubernetes jobs.
langfuse.com/blog/2026-06-22…
It was clear that as soon as we'd launch it people would be all over us. Great to work in an environment that trusts new engineers to work on sth so crucial within their first weeks - but definitely nothing that I wanted to get wrong.
Did this on Langfuse’s CI recently. Told Fable to make our pipelines 30s faster end-to-end, iterating on real CI timings with old runs as the baseline.
10% speed up over night - quite the win for an already very fast pipeline (5min).
Our CI is already fast (~5min) so wins are non-obvious. I Optimizing CI is normally tedious with having 2 wait between attempts. Fable just pushed through: thesis, test run, repeat until it reached the goal.