Comments (8)
I'll post my post processing changes soon.
from metaseq.
I'm particularly interested in how you concatenated messages of the Pushshift dataset (what delimiter did you use e.g.), as I'm trying to get a hand on fine-tuning OPT as a dialogue model.
It's not about how you flattened it, but more about the final format of data.
from metaseq.
Hello, thank you so much for your excellent work. I want to reproduce the training results. Would you mind providing the complete data preprocessing scripts?
from metaseq.
Since the datasets are already publicly available, we aren't hosting and offering the dataset for download.
from metaseq.
I see, thanks for the quick reply.
Section 2.3 of the papaer mentions some filtering & preprocessing done for each of the corpora. Are these scripts released? Where can I find them?
from metaseq.
Great, thanks so much! Looking forward!
from metaseq.
I'm particularly interested in how you concatenated messages of the Pushshift dataset (what delimiter did you use e.g.), as I'm trying to get a hand on fine-tuning OPT as a dialogue model.
It's not about how you flattened it, but more about the final format of data.
One message per line. I wish I had done better.
from metaseq.
https://gist.github.com/stephenroller/8738a3e4fcbeae23dc5dbb87c8745d87 regex processing here
from metaseq.
Related Issues (20)
- How to finetune from a consolidated model ? HOT 1
- Incorrect md5sums after running reshard_fsdp.py on OPT-175B HOT 2
- Converting OPT-175B tokenizer to HF format? HOT 2
- downloading opt-66B part7 get access denied HOT 1
- Confirm md5sums after running reshard_fsdp.py on OPT-175B #702 HOT 3
- Add type hints to all methods
- FSDP is incompatible with BF16 HOT 4
- OPT and LLaMA HOT 1
- load checkpoint failed when training with multi-nodes. HOT 1
- Grammatical Error Correction (GEC) prompt for OPT-IML
- train opt-125M from scratch HOT 1
- Possible feature and bugfix contributions from Microsoft research team's fork of Metaseq HOT 4
- OPT在中文对话上表现如何呢?
- Access request for opt-175b HOT 1
- Process blocks when deploying OPT-1.3B with FasterTransformer
- How can I pretrain an opt-model with the codes?
- setup to pyproject
- Weights/Code for CM3Leon HOT 2
- I change Num_head of OPT-1.3b,and it cause CUDA Error: IndexSelectLargeIndex,
- How to load the checkpoints into a HF model?
Recommend Projects
-
React
A declarative, efficient, and flexible JavaScript library for building user interfaces.
-
Vue.js
🖖 Vue.js is a progressive, incrementally-adoptable JavaScript framework for building UI on the web.
-
Typescript
TypeScript is a superset of JavaScript that compiles to clean JavaScript output.
-
TensorFlow
An Open Source Machine Learning Framework for Everyone
-
Django
The Web framework for perfectionists with deadlines.
-
Laravel
A PHP framework for web artisans
-
D3
Bring data to life with SVG, Canvas and HTML. 📊📈🎉
-
Recommend Topics
-
javascript
JavaScript (JS) is a lightweight interpreted programming language with first-class functions.
-
web
Some thing interesting about web. New door for the world.
-
server
A server is a program made to process requests and deliver data to clients.
-
Machine learning
Machine learning is a way of modeling and interpreting data that allows a piece of software to respond intelligently.
-
Visualization
Some thing interesting about visualization, use data art
-
Game
Some thing interesting about game, make everyone happy.
Recommend Org
-
Facebook
We are working to build community through open source technology. NB: members must have two-factor auth.
-
Microsoft
Open source projects and samples from Microsoft.
-
Google
Google ❤️ Open Source for everyone.
-
Alibaba
Alibaba Open Source for everyone
-
D3
Data-Driven Documents codes.
-
Tencent
China tencent open source team.
from metaseq.