Comments (3)
It's a notation from LoRA. LoRA uses the product of two matrices BA to create a low-rank matrix. In the original LoRA design, BA should be added to W. Multi means instead of adding to W, we use elementwise multiplication. In fact, if we (i) change add to multiply, (ii) set the rank is set to 1, and (iii) only keep the three internal parameters, we can derive IA3 from LoRA.
from t-few.
Hi @HaokunLiu , is there any work that shows the equation of this method? What is the meaning of elementwise multiplication? Thanks!
from t-few.
Wait, somehow I didn't see your reply. Sorry.
For the equation of LoRA, you can refer to their paper. https://arxiv.org/abs/2106.09685
Element-wise multiplication simply means when you have two tensors of the same shape (or one tensor is expandable to have the same shape as the other), for instance X = (5, 3, 16) 3D-tensor, Y = (3, 16) 2D-matrix. you do multiplication at every location. So the output Z = (5, 3, 16) 3D-tensor will be Z_(i,j,k) = X(i,j,k) * Y(j,k), for all the i in {0,1,...,3}, j in {0, 1,2}, and k in {0, 1, ..., 16).
from t-few.
Related Issues (20)
- What is the meaning of score_gt and score_cand? HOT 6
- Accuracy could not match with the log when load_model HOT 10
- Validation score on WSC decreases with training HOT 3
- Sum of logprobs in the probability space adds up to values above 1 HOT 2
- AttributeError: 'DistributedDataParallel' object has no attribute 'save_checkpoint' HOT 1
- save dev_pred.txt and test_pred.txt for RTE and ANLI HOT 2
- How is l_ff created? HOT 1
- How long it will take for pretraining the model using A100(80G)? HOT 2
- Where are the loss function changes in the codebase? HOT 1
- IA3 implementation doesn't add parameters for feedforward layers HOT 5
- questions from your paper HOT 1
- results for LoRA HOT 1
- t-few for decoder only models
- Multi-task batching HOT 4
- question about intrinsic.py HOT 1
- Creation of the `decoder_attention_mask` while evaluating HOT 1
- Could your please give a detailed explanation for the "rank classification"? HOT 1
- Make use of the model, datasets and strategy to classify sentences as urgent not urgent HOT 1
- Issue on the install of first experiment and deepspeed in Windows HOT 1
Recommend Projects
-
React
A declarative, efficient, and flexible JavaScript library for building user interfaces.
-
Vue.js
🖖 Vue.js is a progressive, incrementally-adoptable JavaScript framework for building UI on the web.
-
Typescript
TypeScript is a superset of JavaScript that compiles to clean JavaScript output.
-
TensorFlow
An Open Source Machine Learning Framework for Everyone
-
Django
The Web framework for perfectionists with deadlines.
-
Laravel
A PHP framework for web artisans
-
D3
Bring data to life with SVG, Canvas and HTML. 📊📈🎉
-
Recommend Topics
-
javascript
JavaScript (JS) is a lightweight interpreted programming language with first-class functions.
-
web
Some thing interesting about web. New door for the world.
-
server
A server is a program made to process requests and deliver data to clients.
-
Machine learning
Machine learning is a way of modeling and interpreting data that allows a piece of software to respond intelligently.
-
Visualization
Some thing interesting about visualization, use data art
-
Game
Some thing interesting about game, make everyone happy.
Recommend Org
-
Facebook
We are working to build community through open source technology. NB: members must have two-factor auth.
-
Microsoft
Open source projects and samples from Microsoft.
-
Google
Google ❤️ Open Source for everyone.
-
Alibaba
Alibaba Open Source for everyone
-
D3
Data-Driven Documents codes.
-
Tencent
China tencent open source team.
from t-few.