Comments (4)
Hi,
Your command is bound to give errors. You are missing a few things:
--langs hi_IN (since there's no language token for Sanskrit you may have to use the one for Hindi. Since you don't provide a language token the code bugs out as it uses some default language token that mbart tokenizer doesn't recognize and segments it weirdly.)
Don't use fp16 if you use multiple GPUs. Regardless I've had mixed results with fp16 so I'd avoid it.
--batch_size 16 (I think you meant this to be 16 sentences. So you need the flag --batch_size_indicates_lines. By default this means number of tokens in batch.)
Additionally, be careful about learning rates and dropouts etc. Good luck!
from yanmtt.
Hi,
Thanks for your reply.
I added the arguments.
!python pretrain_nmt.py -n 1 -nr 0 -g 1 --use_official_pretrained --langs hi_IN --batch_size_indicates_lines --pretrained_model "facebook/mbart-large-50" --model_path "facebook/mbart-large-50" --tokenizer_name_or_path "facebook/mbart-large-50" --mono_src "/content/yanmtt/cleaned_Sanskrit_text_for_LM.txt" --shard_files --batch_size 1
You see,I've put the batch size as 1 and even them I'm getting this:
RuntimeError: CUDA out of memory. Tried to allocate 978.00 MiB (GPU 0; 15.90 GiB total capacity; 14.69 GiB already allocated; 337.75 MiB free; 14.76 GiB reserved in total by PyTorch)
GPU:
from yanmtt.
Hi,
I can't really help with the GPU memory issue. 16 GBs is a tab bit small. The only thing you can try is limit the maximum sequence length via the --hard_truncate_length option. Find out what's the average sequence length in your corpus and then play with that argument. Btw why not try IndicBart which is much about a third of mbarts size and is more suited for Indic languages? Since its compact you will not run into memory issues.
from yanmtt.
Sure, Thanks.
from yanmtt.
Related Issues (20)
- Improve documentation
- Improve examples
- Binary executables for all python scripts
- CPU support
- Add support for latest version of transformers repo
- Display more information during training
- RuntimeError: The expanded size of the tensor (22) must match the existing size (21) at non-singleton dimension 1. Target sizes: [178, 22, 1] . Tensor sizes: [178, 21, 1]
- Add PEP8 style guide checker workflow
- Add post-norm to the model
- Mixtures of denoisers
- Support all optimizers and schedulers
- Error in BART Monolingual Pre-training. HOT 5
- Evaluation during training BARTforConditionalGeneration pre-training on English corpora HOT 1
- Alternative to installing sentencpiece HOT 11
- Extending IndicBART or IndicBERT HOT 7
- Pretrain Donut model HOT 1
- Problem with __future__ annotation HOT 1
- Could not find the version : tensorflow-gpu==2.3.0
- Disable shared sentencepiece libraries in installation instructions HOT 1
- Error: Invalid new-expression of abstract class type torchdistx::detail::{anonymous}::ProxyVariableHooks
Recommend Projects
-
React
A declarative, efficient, and flexible JavaScript library for building user interfaces.
-
Vue.js
🖖 Vue.js is a progressive, incrementally-adoptable JavaScript framework for building UI on the web.
-
Typescript
TypeScript is a superset of JavaScript that compiles to clean JavaScript output.
-
TensorFlow
An Open Source Machine Learning Framework for Everyone
-
Django
The Web framework for perfectionists with deadlines.
-
Laravel
A PHP framework for web artisans
-
D3
Bring data to life with SVG, Canvas and HTML. 📊📈🎉
-
Recommend Topics
-
javascript
JavaScript (JS) is a lightweight interpreted programming language with first-class functions.
-
web
Some thing interesting about web. New door for the world.
-
server
A server is a program made to process requests and deliver data to clients.
-
Machine learning
Machine learning is a way of modeling and interpreting data that allows a piece of software to respond intelligently.
-
Visualization
Some thing interesting about visualization, use data art
-
Game
Some thing interesting about game, make everyone happy.
Recommend Org
-
Facebook
We are working to build community through open source technology. NB: members must have two-factor auth.
-
Microsoft
Open source projects and samples from Microsoft.
-
Google
Google ❤️ Open Source for everyone.
-
Alibaba
Alibaba Open Source for everyone
-
D3
Data-Driven Documents codes.
-
Tencent
China tencent open source team.
from yanmtt.