@commanderdgr8

Building VapuAI - an all-in-one AI image generator. 24 yrs shipping software, now figuring out AI. Sharing prompts, techniques, and honest build updates daily.

Joined April 2009
Day 86 of building VapuAI: pushed yesterday's 6 bug fixes live, fixed 7 more today. bug fixes are necessary to maintain quality. Now I can freely write few blogs I have planned and add more features.
1
34
Ashish Sheth retweeted
Everyone is fine-tuning LLMs. Almost nobody understands what is actually being updated inside the model. That distinction matters because LoRA, QLoRA, LoRA-FA, VeRA, Delta-LoRA and LoRA+ are usually discussed as if they were small variations of the same method. They are not. Some reduce trainable parameters, some reduce activation memory, some reduce the memory occupied by the frozen model, and some change the way the adapter itself is optimized. I'm actually sharing 6 techniques. 1/ LoRA Suppose a layer contains a pretrained weight matrix W. Full fine-tuning would update W directly. LoRA leaves W frozen and represents the update using two much smaller matrices, A and B, so that ΔW = BA. For a square d × d weight matrix, full fine-tuning has d² parameters available to update. A rank-r LoRA adapter has roughly 2dr trainable parameters instead, where r is normally much smaller than d. This is the basic reason LoRA can adapt very large models without training every parameter in them. 2/ LoRA-FA Standard LoRA trains both A and B. LoRA-FA freezes A and trains only B. There is a useful reason for doing this beyond simply reducing the number of trainable parameters. Computing the gradient for A requires retaining the layer input activation. If A is fixed, that gradient is no longer required, which allows LoRA-FA to reduce activation memory as the LoRA rank grows. 3/ QLoRA QLoRA attacks a different part of the memory problem. LoRA makes the adapter small, but the frozen base model can still occupy tens of gigabytes. QLoRA keeps the base model frozen in 4-bit form and trains LoRA adapters through it. The original work used NF4, double quantization and paged optimizers, and demonstrated fine-tuning a 65B model on a single 48GB GPU. This is an important distinction: QLoRA is not simply "LoRA with smaller adapters." The large memory saving comes from quantizing the frozen base model. 4/ VeRA VeRA reduces the adapter itself further. Instead of learning a separate A and B for every adapted layer, it uses frozen random low-rank matrices that can be shared across layers, while learning much smaller scaling vectors. The low-rank basis is therefore fixed. Training mainly determines how strongly different parts of that basis should contribute. This is why VeRA can use considerably fewer trainable parameters than ordinary LoRA. 5/ Delta-LoRA Ordinary LoRA treats W as fixed throughout training. Delta-LoRA relaxes that constraint. A and B are still trained, but the change in their product from one training step to the next is also used to update W. The base weights can therefore move without maintaining the ordinary gradients and optimizer states that full fine-tuning would require for W. 6/ LoRA+ LoRA+ does not introduce another adapter structure. It keeps W frozen and still trains A and B. Its change is in the optimizer: A and B use different learning rates, with B receiving a larger rate. The motivation is that the two LoRA matrices do not behave identically during optimization, so forcing them to use the same learning rate is not necessarily the best choice. Once these are separated by what they actually change, the family becomes much easier to understand. LoRA reduces the number of weights being trained. LoRA-FA also targets activation memory. QLoRA compresses the frozen base model. VeRA reduces the learned adapter parameters further. Delta-LoRA allows the pretrained weights themselves to evolve through low-rank changes, while LoRA+ keeps the LoRA structure and changes its optimization. That is really what PEFT is about => deciding which parts of a very large model actually need to move during adaptation, and which parts can remain fixed.
46
485
12
2,861
126,278
Jev is still invite only. But looking at mastery guides available in X, seems lot of people have already mastered it.
15
Ashish Sheth retweeted
This is why I am on Twitter, to see posts like this
Solving a Rubik's Cube with graph theory.
195
18,407
98
226,745
6,554,563
Ashish Sheth retweeted
Jev is the future
129
264
151
4,987
1,130,652
Opus 5 is useless for coding. period.
10
Ashish Sheth retweeted
When I was in 8th, I blew a fuse at home while playing with circuit boards. Connected 5v DC motor directly to an electric board. That day my father got scared and scolded me to not do experiments like this anymore. When an electrician came to my house to repair anything, I asked him so many questions out of interest that he faded up many times. But I never stoped experimenting with electronics. Sometimes, I got shocks too ;) When I started my engineering, I used to bunk class and go to electronic lab. Working on projects. In my 2nd year, made a portable ECG device (no readymade modules) at 800 rupees. At that time, my professor tested it thoroughly and impressed with what I made. He told me to take admission into IITs, for Masters. If I really want to make something big. Gave GATE exams two times and failed. - 1st time, I didn't qualified with 0.83 marks. - 2nd time, I got COVID. Had severe health issues. - 3rd time, I emailed to an IIT Madras professor who was working on optics and accepting MS admission without GATE. Didn't accepted my application. After that, I was lost. Didn't know what to do. Tried to apply for jobs. Didn't accepted as companies wanted only sales people at that time. I had just one thing in my mind. - Wanted to make something, execute it in the market and see people using it. While making specialised COVID Ventilators with a company, I got an idea for RespiCOz. That evening, I just called my professor and said with clear mind that I am going to make this. I was so clear that I had just one goal. I want to execute this. During that time, I found that Gujarat Government was funding startups for building things. I immediately incorporated the company and started applying for those grants. That's how Brainiac Healthcare started. And got some government grants (From State and Central Gov both). After 3 years of iteration, testing and approvals, we launched our first version of RespiCOz (in Aug 2024). Early days were difficult. The adoption rate was low. I lost some confidence. But after 5-6 months, we started seeing reference orders. The sales started booming. Today, RespiCOz is present in 3+ countries, have more than 300+ Installations. While scaling RespiCOz, I again lost myself what I am going to build next. RespiCOz could be just a start to a big thing. Not the end. While working on RespiCOz datapoints and reading some research papers, found that we haven't explored respiratory care at its full potential. We can do a lot, if we use intelligence. That's how we got an idea for Zapien (@ZapienLife ). Now, we are on a mission to explore respiratory care and give humanity a new perspective of care. Think, just a breath can diagnose your body. Will you believe it? You will see it becoming a reality in few years. Respiratory Technologies for Medical Devices, Consumer Wellness Products and Pharma. All under one roof, with Intelligence. Now, let's build.
125
341
29
2,283
64,956
Cloudflare is rolling out changes to its AI Bot Controls, where users can hint to bot crawlers that the content is allowed to be searched, but not allowed for AI training. The question is what is the guarantee that Google, Apple, Bing etc will not use your content for training their AI even after this settings.
9
What is the point of government if citizens have to fix these problems themselves ? @narendramodi @BBMPSWMSplComm @bbmpcommr
The footpath near my society at Koramangala has been broken since years! I complained to BBMP, sent email to GBA, raised ticket and chatted with their WA helpline but got no response. Finally, decided to fix it myself. It costed me Rs. 2700 including labor charges and the cement to build a ramp. Parents with stroller and senior citizens on wheelchair can use it conveniently now.
23
OpenArch- One awesome GitHub repo worth checking. This contains PyTorch implementation of model architectures mentioned in LLM Architecture gallary from Sebastian Raschka.
15
Except Docker and Kubernetes, which one of these existed 8 years back ( in 2018/19)
14
Social media platforms are a large stage. Influencers are the performers. Others are just supporting cast .
10
My cable collection! Not sure when I will use many of them!!
1
8
Someday, I will also say, that day is today.
2
Tell me how AI has gaslighted you.
11
A very nice LLM Attention Visualizer. This helps in understanding which input token affects the generation of which output tokens.
1
11