@TDataSciencei
iAccount based inAustria
About this account
- Account based in
- Austria
- Connected via
- Web
Account-level information from X, not a live location or the device used for a specific post.
The world's leading publication for data science and artificial intelligence professionals. Submit an Article ✍️ https://nitter.cf/t.co/57pIMegK1o
We're Global 🌏
Joined October 2016
- Tweets64.3K
- Following1.9K
- Followers251K
- Likes8K
Towards Data Science retweeted
Hi all, I wrote an article on @TDataScience, do check it out: towardsdatascience.com/the-a…
Towards Data Science retweeted
How does text go from words to numbers? 🤔
My latest blog published in @TDataScience explores TF-IDF step by step from text to vectors and finally to text classification.
I also look at its limitations and what happens when we encounter a new word.
towardsdatascience.com/from-…
In his recent guide, Conor O'Sullivan shared practical tips for using AI while working on a PhD thesis — but these insights can be equally applied to many other complex, long-term projects involving massive amounts of data and references. towardsdatascience.com/4-way…
In a new deep dive, Tigran Hayrapetyan introduces Guided Merge Sort, an optimized sorting approach that chooses the best options among ordinary and multi-way merge-sort algorithms. towardsdatascience.com/guide…
What if you could transform an open-source LLM into your own, Jev-like single-pass text classifier? Anubhab Banerjee outlines an approach that can accomplish this result. towardsdatascience.com/how-t…
An HR assistant sandbox that wraps a RAG agent in the checks a production system needs: injection filters, ACL pre-filtering, and human approval on high-risk actions.
By Partha Sarkar
towardsdatascience.com/shipa…
"At some point, with no new data and nothing visibly different about the setup, the model's performance on the unseen questions jumped from barely better than random to nearly perfect."
Utkarsh Mangal unpacks the phenomenon of grokking in machine learning. towardsdatascience.com/the-a…
Learn how you can detect hidden shifts in feature relationships to avoid data drift — @BenjaminNweke11 leverages adversarial validation and scikit-learn in his latest tutorial. towardsdatascience.com/how-t…
Can you leverage the power of Jev in the context of scalable knowledge graphs? Partha Sarkar explains what that looks like in practice in his excellent hands-on guide. towardsdatascience.com/graph…
"Cursor, Claude Code and Copilot all sit in front of the same gap, whatever each of them indexes, because the information is not in any repository they can see."
Yonatan Sason zooms in on a structural issue affecting AI-assisted software development. towardsdatascience.com/good-…
AI-generated content has entered training data — and at a massive rate. Abdullahi Dattijo presents his findings from testing three approaches to spot it. towardsdatascience.com/ai-sl…
How do paragraphs "work" as a meaningful unit of content in the context of LLMs? Shuyang Xiang offers an accessible explainer on an often-overlooked topic. towardsdatascience.com/your-…
Not so long ago, building an internal tool at work would take weeks, if not months. @EivindKjos shows how you can cut down the timeline to mere hours with the help of Claude Code. towardsdatascience.com/build…
From stale, contradictory documents to a table divided across a page boundary, Sara Nobrega walks us through some of the adversarial test tests she recommends you run before giving users access to your RAG pipeline. towardsdatascience.com/break…
In a new hands-on guide, Spyros Georgopoulos walks us through five essential scikit-learn defaults you should remember to vet when reviewing AI-generated code. towardsdatascience.com/your-…
"Unlike physical assets, data, more specifically operational data, doesn't have a direct value. But we can still use valuation principles to value data assets."
@Thuwarakesh shares helpful context on the recent news that Googled offered $10m for the data of an airline going through bankruptcy. towardsdatascience.com/googl…
A financial reporting co-pilot for insurance regulatory closes, where deterministic code makes every number and the model supplies the language around them.
By Ari Joury.
towardsdatascience.com/shipa…
What does the GPT-6 Astra release reveal about the industry's ability — or willingness — to measure risk? @BenjaminNweke11 discusses the model's assigned cybersecurity risk level and its implications. towardsdatascience.com/gpt-6…
"This new model looks particularly well suited to many everyday use cases, such as classification (e.g. topic modelling for NPS comments) or LLM-as-a-judge tasks. So, naturally, I decided I had to try it out."
Mariya Mansurova takes Jev, TypeSafe's One System model, for a spin. towardsdatascience.com/a-new…
"AI handed us spare capacity, and our first instinct was to fill it with more AI, instead of asking what the spare capacity was actually for."
Gursimar Singh shares an incisive reflection on the unexpected effects of leaning into agentic AI-powered workflows. towardsdatascience.com/ai-ma…