@DatabendLabsi
iAccount based inJapan!
About this account
- Account based in
- Japan
- Connected via
- Singapore App Store
! X says this location may be affected by a proxy or VPN.
Account-level information from X, not a live location or the device used for a specific post.
Agent-ready Warehouse for analytics, search, AI, and Sandbox. Databend Cloud: https://nitter.cf/t.co/P3nJRYhion GitHub: https://nitter.cf/t.co/0gHU5sln82
SF
Joined September 2021
- Tweets809
- Following133
- Followers2.2K
- Likes386
Pinned Tweet
Millions of context tokens. Thousands of tool calls. TB-scale Trace writes per hour.
See how a leading foundation model company built an end-to-end Agent Trace pipeline on Databend Cloud—turning raw execution data into eval-ready datasets.
Read the customer story 👇
bit.ly/4z19ujk
Same PASS. Very different execution paths.
Jev evaluates each span for relevance and progress, then writes the results to Databend Cloud for SQL analysis. In our runs, evaluating 1M spans would cost about $80.
Read the engineering breakdown: bit.ly/4yfaOyj
Keyword rules miss the meaning in customer feedback 🔍.
We connected Jev to Databend's Python UDF server so SQL can filter, classify, and rank records using model probabilities 📊.
Databend runs the query; Jev makes the bounded decision ✅.
Explore: bit.ly/4xFq4Dp
Cloud Agent is in Preview 🚀 — your AI assistant for Databend Cloud.
Ask in natural language 💬.
Analyze data and workloads 📊.
Take action with your approval ✅.
One example: “Where can I reduce warehouse costs?”
bit.ly/4xXfQix
🔥We’re launching the Agent Trace Solution.
Capture complete agent runs, preserve evolving JSON, and turn new spans into query-ready and eval-ready data—on one S3-backed lakehouse.
Built for production scale 🔍 bit.ly/3UIiXMJ
Async refresh shouldn’t mean stale answers.
Databend materialized views use Change Tracking + checkpoints, then choose Fresh, Hybrid read fix, or Live Fallback at query time.
Refresh lag changes query cost—not correctness.
Deep dive 👉 bit.ly/4xav8Q7
One lakehouse. Five production scenarios.
Agent traces, AI apps, industrial telemetry, public-sector data, and financial archives—built on Databend Cloud with the same pattern: preserve raw data, process changes incrementally, isolate compute.
Read 👉 bit.ly/4gDTQlF
On a controlled 100M-row test:
62.9% → 13.6% of blocks scanned.
63.5M → 13.6M rows read.
The difference: a workload-driven Databend Cluster Key.
Learn how to choose columns, key order & granularity—and verify the result 👉 bit.ly/4y3VF2k
Recent Snowflake threads repeat the warning: clustering helps pruning, but can burn credits when keys miss the workload.
Cluster keys aren’t indexes. Databend pairs automatic clustering with block metrics + SQL RECLUSTER controls.
Deep dive → bit.ly/4xSfQjz
DeepSeek Harness puts Agent Trace in focus: long contexts, tool calls, and retries create a data problem.
Production proof: a leading AI company uses Databend Cloud for TB-scale ingestion, turning spans into eval-ready data—with less OSS overhead.
Reed 👉 bit.ly/4z19ujk
Modern apps & AI workloads need more than fast SQL.
Databend Cloud brings together:
→ Cloud-native lakehouse
→ Real-time ELT
→ Multimodal analytics in one SQL engine
→ Governance, isolation & cost control
Start with free credits 👇
app.databend.com/login
Agent Trace pipelines fail at the boundaries: schema drift, small writes, and early offset commits.
bend-ingest-kafka batches events, commits after COPY succeeds, then Databend Stream + MERGE processes new data incrementally.
Engineering deep dive 👉 bit.ly/4g6pDLH
June–July at Databend: 182 updates across 15 nightlies + 2 patches.
Smarter optimizer stats. R-Tree Spatial Index Join. Paimon Catalog. More reliable Stream + Task pipelines for model Eval and Agent Trace workloads.
Read the engineering update 👉 bit.ly/3TDO5wj
Snowflake teams keep revisiting the FinOps loop: right-size warehouses, tune auto-suspend, audit queries, assign cost owners.
Databend Cloud starts simpler: transparent pay-as-you-use compute, open Parquet, 90%+ Snowflake SQL parity, and BYOC.
Compare 👉 bit.ly/4bnDz2h
Global sorts make optimizer stats accurate and expensive.
Databend combines KLL for one-pass quantiles, Top-N for exact hot values, and CMS for a wider hot set.
In our benchmark, CMS cut equality q-error from 25.354 to 1.355. 🚀
Deep dive: bit.ly/4pKP3my
Filtering a window function shouldn’t require another query layer.
QUALIFY keeps ranking and filtering together for deduplication, latest rows, and Top N per group.
Cleaner SQL. Faster reviews.
See 6 practical patterns in Databend 👉bit.ly/4yIMLZh
#DataEngineering #SQL
When agent traces reach TB scale, is “observability” still the right mental model?
They’re also eval datasets, replay records, and training feedback.
Should teams build the underlying data infrastructure themselves, or use a platform like Databend?
Here’s a deep dive into the engineering considerations👉 bit.ly/4v7Ddno
Immersive Translate needed product analytics without a heavy big data stack.
So they kept it simple:
📥 NDJSON on S3
⚙️ Databend Cloud scheduled Tasks
📊 SQL analytics + BI dashboards
💰 elastic warehouses with auto-suspend
🚀 POC completed in one afternoon.
Read the customer story: bit.ly/4bw6EZn
#DataEngineering #CloudDataWarehouse
AI can generate an AST Visitor migration fast.
Proving it didn’t change SQL semantics is the hard part.
Inside Databend: source-derived AST metadata, runtime traces, independent comparison, real SQL corpora, and coverage feedback.
Read the blog👉 bit.ly/4wGL6RX
#Database #AI