@commanderdgr8i
iAccount based inIndia
About this account
- Account based in
- India
- Connected via
- India App Store
Account-level information from X, not a live location or the device used for a specific post.
Building VapuAI - an all-in-one AI image generator. 24 yrs shipping software, now figuring out AI. Sharing prompts, techniques, and honest build updates daily.
Joined April 2009
- Tweets0
- Following0
- Followers0
- Likes0
Day 86 of building VapuAI: pushed yesterday's 6 bug fixes live, fixed 7 more today.
bug fixes are necessary to maintain quality.
Now I can freely write few blogs I have planned and add more features.
We have seen wrappers around models. This one is wrapper around your harness!!
github.com/mvschwarz/openrig
Ashish Sheth retweeted
Everyone is fine-tuning LLMs. Almost nobody understands what is actually being updated inside the model.
That distinction matters because LoRA, QLoRA, LoRA-FA, VeRA, Delta-LoRA and LoRA+ are usually discussed as if they were small variations of the same method. They are not. Some reduce trainable parameters, some reduce activation memory, some reduce the memory occupied by the frozen model, and some change the way the adapter itself is optimized.
I'm actually sharing 6 techniques.
1/ LoRA
Suppose a layer contains a pretrained weight matrix W. Full fine-tuning would update W directly. LoRA leaves W frozen and represents the update using two much smaller matrices, A and B, so that ΔW = BA.
For a square d × d weight matrix, full fine-tuning has d² parameters available to update. A rank-r LoRA adapter has roughly 2dr trainable parameters instead, where r is normally much smaller than d. This is the basic reason LoRA can adapt very large models without training every parameter in them.
2/ LoRA-FA
Standard LoRA trains both A and B. LoRA-FA freezes A and trains only B.
There is a useful reason for doing this beyond simply reducing the number of trainable parameters. Computing the gradient for A requires retaining the layer input activation. If A is fixed, that gradient is no longer required, which allows LoRA-FA to reduce activation memory as the LoRA rank grows.
3/ QLoRA
QLoRA attacks a different part of the memory problem. LoRA makes the adapter small, but the frozen base model can still occupy tens of gigabytes.
QLoRA keeps the base model frozen in 4-bit form and trains LoRA adapters through it. The original work used NF4, double quantization and paged optimizers, and demonstrated fine-tuning a 65B model on a single 48GB GPU.
This is an important distinction: QLoRA is not simply "LoRA with smaller adapters." The large memory saving comes from quantizing the frozen base model.
4/ VeRA
VeRA reduces the adapter itself further. Instead of learning a separate A and B for every adapted layer, it uses frozen random low-rank matrices that can be shared across layers, while learning much smaller scaling vectors.
The low-rank basis is therefore fixed. Training mainly determines how strongly different parts of that basis should contribute. This is why VeRA can use considerably fewer trainable parameters than ordinary LoRA.
5/ Delta-LoRA
Ordinary LoRA treats W as fixed throughout training. Delta-LoRA relaxes that constraint.
A and B are still trained, but the change in their product from one training step to the next is also used to update W. The base weights can therefore move without maintaining the ordinary gradients and optimizer states that full fine-tuning would require for W.
6/ LoRA+
LoRA+ does not introduce another adapter structure. It keeps W frozen and still trains A and B.
Its change is in the optimizer: A and B use different learning rates, with B receiving a larger rate. The motivation is that the two LoRA matrices do not behave identically during optimization, so forcing them to use the same learning rate is not necessarily the best choice.
Once these are separated by what they actually change, the family becomes much easier to understand.
LoRA reduces the number of weights being trained. LoRA-FA also targets activation memory. QLoRA compresses the frozen base model. VeRA reduces the learned adapter parameters further. Delta-LoRA allows the pretrained weights themselves to evolve through low-rank changes, while LoRA+ keeps the LoRA structure and changes its optimization.
That is really what PEFT is about => deciding which parts of a very large model actually need to move during adaptation, and which parts can remain fixed.
Jev is still invite only. But looking at mastery guides available in X, seems lot of people have already mastered it.
How GLM built its own infrastructure
z.ai/blog/glm-built-its-infe…
Introducing Bonsai 2 27B.
prismml.com/news/bonsai-2-27…
Ashish Sheth retweeted
When I was in 8th, I blew a fuse at home while playing with circuit boards.
Connected 5v DC motor directly to an electric board.
That day my father got scared and scolded me to not do experiments like this anymore.
When an electrician came to my house to repair anything, I asked him so many questions out of interest that he faded up many times.
But I never stoped experimenting with electronics.
Sometimes, I got shocks too ;)
When I started my engineering, I used to bunk class and go to electronic lab.
Working on projects.
In my 2nd year, made a portable ECG device (no readymade modules) at 800 rupees.
At that time, my professor tested it thoroughly and impressed with what I made.
He told me to take admission into IITs, for Masters. If I really want to make something big.
Gave GATE exams two times and failed.
- 1st time, I didn't qualified with 0.83 marks.
- 2nd time, I got COVID. Had severe health issues.
- 3rd time, I emailed to an IIT Madras professor who was working on optics and accepting MS admission without GATE. Didn't accepted my application.
After that, I was lost. Didn't know what to do.
Tried to apply for jobs. Didn't accepted as companies wanted only sales people at that time.
I had just one thing in my mind.
- Wanted to make something, execute it in the market and see people using it.
While making specialised COVID Ventilators with a company, I got an idea for RespiCOz.
That evening, I just called my professor and said with clear mind that I am going to make this.
I was so clear that I had just one goal. I want to execute this.
During that time, I found that Gujarat Government was funding startups for building things.
I immediately incorporated the company and started applying for those grants.
That's how Brainiac Healthcare started.
And got some government grants (From State and Central Gov both).
After 3 years of iteration, testing and approvals, we launched our first version of RespiCOz (in Aug 2024).
Early days were difficult. The adoption rate was low. I lost some confidence.
But after 5-6 months, we started seeing reference orders. The sales started booming.
Today, RespiCOz is present in 3+ countries, have more than 300+ Installations.
While scaling RespiCOz, I again lost myself what I am going to build next.
RespiCOz could be just a start to a big thing. Not the end.
While working on RespiCOz datapoints and reading some research papers, found that we haven't explored respiratory care at its full potential.
We can do a lot, if we use intelligence.
That's how we got an idea for Zapien (@ZapienLife ).
Now, we are on a mission to explore respiratory care and give humanity a new perspective of care.
Think, just a breath can diagnose your body.
Will you believe it?
You will see it becoming a reality in few years.
Respiratory Technologies for Medical Devices, Consumer Wellness Products and Pharma.
All under one roof, with Intelligence.
Now, let's build.
Cloudflare is rolling out changes to its AI Bot Controls, where users can hint to bot crawlers that the content is allowed to be searched, but not allowed for AI training.
The question is what is the guarantee that Google, Apple, Bing etc will not use your content for training their AI even after this settings.
What is the point of government if citizens have to fix these problems themselves ? @narendramodi @BBMPSWMSplComm @bbmpcommr
The footpath near my society at Koramangala has been broken since years!
I complained to BBMP, sent email to GBA, raised ticket and chatted with their WA helpline but got no response.
Finally, decided to fix it myself.
It costed me Rs. 2700 including labor charges and the cement to build a ramp.
Parents with stroller and senior citizens on wheelchair can use it conveniently now.
OpenArch- One awesome GitHub repo worth checking.
This contains PyTorch implementation of model architectures mentioned in LLM Architecture gallary from Sebastian Raschka.
Social media platforms are a large stage. Influencers are the performers. Others are just supporting cast .
Nuances you should be aware while using OpenRouter.
mmoustafa.com/blog/so-you-wa…