@newfront

OSS Engineer | Speaker • Trainer | #DatabricksMVP | Author @OReillyMed | ❤️ #ApacheSpark. ❤️ #Dogs. #DatabricksMVP. Views are my own

California, USA
Joined November 2008
Scott Haines retweeted
Join us next Wednesday for Schema-on-Read Without Regrets! Geir Alstad (Gabler AS) and @newfront (@databricks) will walk through Spark Variant shredding in a real pipeline. 🗓️ Wed, Sept 30 · 8:00 AM PT 🎟️ Register: luma.com/open-z92a #ApacheSpark
1
1
12
1,464
Scott Haines retweeted
How well do you actually know Protocol Buffers? We put together a quick interactive quiz covering schema design, field rules, and wire format mechanics. protobuf.com/quiz/ The questions are now randomly selected from a question bank so retrying will give you a few new questions. #protobuf #grpc #quiz
1
2
12
768
Scott Haines retweeted
Omnigent v0.14.0 is now available! 🎉 This release makes it easier to work across GitHub repositories, configure sessions, and keep track of ongoing work. 🔀 Multi-repo GitHub PRs and sandboxes 🎛️ Redesigned composer 🪟 Claude SDK /compact 🗄️ Gensee, Kubernetes, CockroachDB 🔁 Approvals that survive restarts Learn more 👉 omnigent.ai/releases/0.14.0 #Omnigent #AIAgents #OpenSource
4
7
1
27
3,096
Scott Haines retweeted
At Airflow Summit, Meni Shmueli from DataFlint covers how #ApacheSpark and @ApacheAirflow run together in production, including splitting CPU work (like reading Parquet) off GPU machines so AI jobs aren’t paying GPU rates for CPU-only steps. Apache Spark 4.2 is out: auto CDC and Real-Time Mode for streaming, DataSketches as first-class operators, and Spark Connect so client/server queries come back in milliseconds instead of spinning up a full app. #SparkConnect #DataSketches #ApacheAirflow
1
1
12
1,517
RT @ApacheSpark: 📣 Schema-on-Read Without Regrets: Lessons from Apache Spark™ Variant shredding in Production Join Geir E. Alstad (Gabler…
1
4
Scott Haines retweeted
The Replit-Databricks integration is now generally available, with new native Lakebase support! Now, you can use @Replit to build full-stack Databricks Apps on live, governed enterprise data. Replit Agent can automatically provision a Lakebase database, while AI-proposed schema changes require human approval. Build, deploy, and iterate in one flow. See it in action: replit.com/blog/databricks20…
8
13
2
88
12,299
Scott Haines retweeted
What’s New in MLflow: September 2026 Roundup nitter.cf/i/broadcasts/1kKzDPqQY…
1
6
483
Scott Haines retweeted
3 weeks out: Open Lakehouse + AI Meetup is coming to Paris! 🇫🇷 Agenda: 🔹 @unitycatalog_io (You Can Go Your Own Way....Go Your Own Way) 🔹 @omnigent_ai — A Meta Harness for AI Agents 🔹 AI Agents at Every Scale: Personal, Team, and Data Workflows Then networking, light bites & swag. 📅 Wed, Sept 23 | 6:00–9:30 PM CEST 📍 La Fondation | Paris, FR 🎟️Register: usergroups.databricks.com/ev… #OpenLakehouse #AI
1
4
7
809
Scott Haines retweeted
Delta Lake 4.4.0 is out! 🙌 It simplifies engine integrations for catalog-managed tables via Delta Kernel and the UC Delta APIs. Delta Spark now defaults to Spark 4.2. Identity columns in CREATE TABLE. Flink primary-key upserts. 🔗 Learn more: delta.io/blog/2026-08-20-sim… #DeltaLake #ApacheSpark #OpenSource
3
19
1,590
Scott Haines retweeted
Apache Spark is built for large, fault-tolerant jobs: query plans, stages, tasks, and retries. On small local data, such as about 20 MB of Parquet or JSON, those scheduling steps can add roughly 100 milliseconds each, and the iteration loop gets slow. Project Feather is a SPIP to make that laptop and local-mode loop faster without a new engine or API changes. Spark committer Daniel Tenedorio, a co-author of the proposal, walks through the work here 👉 youtu.be/p9ABo1BK6Zc #ApacheSpark #OpenSource #DataEngineering #Spark
4
32
2,565
Scott Haines retweeted
Mark your calendars! 🗓 The Delta Lake Community Meetup is back on Tuesday, September 8 at 9:00 AM PT. Here’s what we’re covering: 🔹Delection Vectors Support in CDF using Kernel – Stephen Carman (@Databricks) breaks down technical updates & key challenge fixes 🔹delta-rs 1.0 – Tyler Croy (Buoyant Data) highlights the latest features and major updates RSVP 👇 luma.com/delta-7kbm #deltalake #opensource #oss
1
10
616
Scott Haines retweeted
DevConnect is back on tour for H2 2026, built for experienced Databricks builders who want a closer look at what’s worth shipping next. Expect deep dives into: • Lakehouse//RT, LTAP, Disaster Recovery, Multi Table Transactions, Catalog Managed Commits, and open table interoperability • MCPs, Unity AI Gateway, ZeroOps, and Omnigent to cut boilerplate, reduce context switching, and keep costs in check • Apps, Genie, and custom chat dashboards to put more builds in the hands of non-data teammates Find a stop near you and register: bit.ly/databricks-devconnect
2
5
36
4,318
Scott Haines retweeted
1 week away ⏳ delta-explain: Making Delta Lake File Pruning Visible, and Testable in CI See how Delta Lake file pruning is attributed to partition pruning and data skipping from transaction-log metadata, then turn that analysis into a CI contract before regressions hit production. 📅 Aug 18 | 9AM PT (virtual) 🎟️ Register: luma.com/delta-odhl #DeltaLake #OpenSource
1
1
4
772
Scott Haines retweeted
Apache Spark 4.2 is now available and it continues maturing Spark Connect. 🙌 It's a thin client over gRPC and Arrow: the client builds the plan, the server runs it, results return as Arrow batches, no full runtime or JVM on the client. 4.2 also keeps closing the compatibility gap with Spark Classic. 🔗Download Spark 4.2: spark.apache.org/downloads.h… #ApacheSpark #SparkConnect #PySpark #ApacheArrow
3
16
1,698
This is going to be an exciting evening (not to mention fall in the Bellevue area), chock full of wonderful presentations on @ApacheSpark, @DeltaLakeOSS , and @ApacheIceberg, with a panel discussion on the entire #openlakehouse ecosystem. Sign up early, you won’t want to miss this.
📣 Join us for the next Open Lakehouse + AI Meetup! On the agenda: 🔹 Modern Analytics at GPU Speed: Accelerating Spark, Delta Lake, Iceberg, and Spark Connect 🔹 A Unified Future for Delta Lake and Apache Iceberg 🔹 Panel: Connecting the Open Lakehouse and AI Ecosystem 🔹 Networking with light bites 🗓️ Wed, October 7 📍 Bellevue | @databricks office 🎟️ Register: luma.com/open-qljr #OpenLakehouse #DeltaLake #ApacheIceberg #ApacheSpark
2
273
Scott Haines retweeted
Replying to @databricks
@databricks has the fastest and lowest latency on Kimi K3 (max)! Great job by Databricks AI team! artificialanalysis.ai/models…
17
249
8
749
26,801
Scott Haines retweeted
New video: @lisancao + Szehon Ho (Spark committer / Iceberg PMC) on Data Source V2 (DSV2) and how Spark connects to Iceberg and Delta Lake. Inside the episode: 🔧 What DSV2 is: richer metadata beyond Hive Metastore 🔄 Unified DML across Iceberg and Delta Lake 📚 Pluggable catalogs + what’s next in Spark 4 Full conversation below 👇 youtu.be/MOQ3G8Ul4Gk #ApacheSpark #DSV2 #ApacheIceberg #DeltaLake
2
14
1,700
Scott Haines retweeted
Full house at the first ever inaugural Rust AI Begins meetup in SF AWS Loft 🔥. With an amazing community pushing the boundaries of what’s possible with Rust and AI. We are excited to share our work about high permanence Rust Engine for fresh context @LinghuaJ along with infra friends at @databricks , @temporalio , @valkey_io , @LakeSail , @ApacheIggy / laserdata and many amazing Rust projects that defines today's landscape for Rust AI Infrastructure. We’re also thrilled to see such passionate and in-depth discussions from the community. Thanks a lot for hosting the amazing mini conf @ChiefScientist @newfront and we are excited to build along with the amazing Rust community 🩷 !
3
8
1
25
7,614