A Survey on Data Synthesis and Augmentation for Large Language Models

December 19, 2024 · View on GitHub

A collection of AWESOME papers on data synthesis and augmentation for Large Language Models.

🌏Please check out our survey paper: https://arxiv.org/abs/2410.12896.

Table of Contents

Taxonomy

Data Augmentation

PaperPublished inCode/Project
Open-source large language models outperform crowd workers and approach ChatGPT in text-annotation tasksarxiv 2023-
ChatGPT outperforms crowd workers for text-annotation tasksPNAS 2023-
Q: How to Specialize Large Vision-Language Models to Data-Scarce VQA Tasks? A: Self-Train on Unlabeled Images!CVPR 2023https://github.com/codezakh/SelTDA
Mind's eye: Grounded language model reasoning through simulationarxiv 2022-
Chatgpt-4 outperforms experts and crowd workers in annotating political twitter messages with zero-shot learningarxiv 2023-
Can chatgpt reproduce human-generated labels? a study of social computing tasksarxiv 2023-
CORE: A retrieve-then-edit framework for counterfactual data generationEMNLP 2022https://github.com/tanay2001/CORE
Diversify your vision datasets with automatic diffusion-based augmentationNeurlPS 2023https://github.com/lisadunlap/ALIA
Llamax: Scaling linguistic horizons of llm by enhancing translation capabilities beyond 100 languagesEMNLP 2024https://github.com/CONE-MT/LLaMAX/
Gpt3mix: Leveraging large-scale language models for text augmentationEMNLP 2021https://github.com/naver-ai/hypermix
Closing the loop: Testing chatgpt to generate model explanations to improve human labelling of sponsored content on social mediaxAI 2023https://github.com/thalesbertaglia/chatgpt-explanations-sponsored-content/
Data augmentation using llms: Data perspectives, learning paradigms and challengesarxiv 2024-
Coannotating: Uncertainty-guided work allocation between human and large language models for data annotationsEMNLP 2023https://github.com/SALT-NLP/CoAnnotating

Data Synthesis

PaperPublished inCode/Project
AlpaGasus: Training A Better Alpaca with Fewer Dataarxiv 2023https://lichang-chen.github.io/AlpaGasus/
TinyStories: How Small Can Language Models Be and Still Speak Coherent English?arxiv 2023-
Textbooks Are All You Need II: phi-1.5 technical reportarxiv 2023-
Multimodal Self-Instruct: Synthetic Abstract Image and Visual Reasoning Instruction Using Language Modelarxiv 2023https://multi-modal-self-instruct.github.io/
Genixer: Empowering Multimodal Large Language Models as a Powerful Data Generatorarxiv 2023https://github.com/zhaohengyuan1/Genixer
Solving Quantitative Reasoning Problems with Language Modelsarxiv 2022-
WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instructarxiv 2023-
WizardCoder: Empowering Code Large Language Models with Evol-Instructarxiv 2023https://github.com/nlpxucan/WizardLM.
Magicoder: Empowering Code Generation with OSS-Instructarxiv 2023-
VILA2^2: VILA Augmented VILAarxiv 2024-
Rephrasing the Web: A Recipe for Compute and Data-Efficient Language Modelingarxiv 2024-
Self-Instruct: Aligning Language Models with Self-Generated InstructionsACL Anthology 2023https://github.com/yizhongw/self-instruct
STaR: Bootstrapping Reasoning With Reasoningarxiv 2022-

Full Lifecycle of LLM

Data preparation

PaperPublished inCode/Project
Tinystories: How small can language models be and still speak coherent english?arxiv 2023https://huggingface.co/roneneldan
Controllable dialogue simulation with in-context learningarxiv 2022https://github.com/Leezekun/dialogic
Genie: Achieving human parity in content-grounded datasets generationarxiv 2024-
Case2Code: Learning Inductive Reasoning with Synthetic Dataarxiv 2024https://github.com/choosewhatulike/case2code
Magicoder: Empowering Code Generation with OSS-Instruct41 ICMLhttps://github.com/ise-uiuc/magicoder
Self-Instruct: Aligning Language Models with Self-Generated Instructionsarxiv 2023https://arxiv.org/abs/2212.10560
Wizardlm: Empowering large language models to follow complex instructionsarxiv 2023https://github.com/nlpxucan/WizardLM
Augmenting Math Word Problems via Iterative Question Composingarxiv 2024https://huggingface.co/datasets/Vivacem/MMIQC
Common 7b language models already possess strong math capabilitiesarxiv 2024https://github.com/Xwin-LM/Xwin-LM
Mammoth: Building math generalist models through hybrid instruction tuningarxiv 2023https://tiger-ai-lab.github.io/MAmmoTH/
Enhancing chat language models by scaling high-quality instructional conversationsarxiv 2024https://github.com/thunlp/UltraChat
Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothingarxiv 2024https://magpie-align.github.io/
GenQA: Generating Millions of Instructions from a Handful of Promptsarxiv 2024https://huggingface.co/datasets/tomg-group-umd/GenQA
Sharegpt4v: Improving large multi-modal models with better captionsarxiv 2023https://sharegpt4v.github.io/
What makes for good visual instructions? synthesizing complex visual reasoning instructions for visual instruction tuningarxiv 2023https://github.com/RUCAIBox/ComVint
Stablellava: Enhanced visual instruction tuning with synthesized image-dialogue dataarxiv 2023https://github.com/icoz69/StableLLAVA
Anygpt: Unified multimodal llm with discrete sequence modelingarxiv 2024https://junzhan2000.github.io/AnyGPT.github.io/
Multimodal Self-Instruct: Synthetic Abstract Image and Visual Reasoning Instruction Using Language Modelarxiv 2024https://github.com/zwq2018/Multi-modal-Self-instruct
Chartllama: A multimodal llm for chart understanding and generationarxiv 2023https://tingxueronghua.github.io/ChartLlama/
Genixer: Empowering Multimodal Large Language Models as a Powerful Data Generatorarxiv 2023https://github.com/zhaohengyuan1/Genixer
Open-Source LLMs for Text Annotation: A Practical Guide for Model Setting and Fine-Tuningarxiv 2024https://osf.io/ctgqx/
ChatGPT Outperforms Crowd-Workers for Text-Annotation TasksNAS 2023-
Can Large Language Models Aid in Annotating Speech Emotional Data? Uncovering New Frontiersarxiv 2023-
Can ChatGPT Reproduce Human-Generated Labels? A Study of Social Computing Tasksarxiv 2023-
Chatgpt-4 outperforms experts and crowd workers in annotating political twitter messages with zero-shot learningarxiv 2023-
Unraveling chatgpt: A critical analysis of ai-generated goal-oriented dialogues and annotationsICIAAI-
FullAnno: A Data Engine for Enhancing Image Comprehension of MLLMsarxiv 2024https://arcana-project-page.github.io/
DISCO: Distilling counterfactuals with large language modelsarxiv 2023https://github.com/eric11eca/disco
Tinygsm: achieving> 80% on gsm8k with small language modelsarxiv 2023-
Gpt3mix: Leveraging large-scale language models for text augmentationarxiv 2021https://github.com/naver-ai/hypermix
CORE: A retrieve-then-edit framework for counterfactual data generationarxiv 2022https://github.com/tanay2001/CORE
Diversify your vision datasets with automatic diffusion-based augmentationarxiv 2023https://github.com/lisadunlap/ALIA
Closing the loop: Testing chatgpt to generate model explanations to improve human labelling of sponsored content on social mediaarxiv 2023https://github.com/thalesbertaglia/chatgpt-explanations-sponsored-content/
Toolcoder: Teach code generation models to use api search toolsarxiv 2023-
Coannotating: Uncertainty-guided work allocation between human and large language models for data annotationarxiv 2023https://github.com/SALT-NLP/CoAnnotating
Does Collaborative Human-LM Dialogue Generation Help Information Extraction from Human Dialogues?arxiv 2023https://boru-roylu.github.io/DialGen
Measuring mathematical problem solving with the math datasetarxiv 2021-
Llemma: An open language model for mathematicsarxiv 2023https://github.com/EleutherAI/math-lm
Code Less, Align More: Efficient LLM Fine-tuning for Code Generation with Data Pruningarxiv 2024-

Pretraining

PaperPublished inCode/Project
VILA2: VILA Augmented VILAarxiv 2024https://github.com/NVlabs/VILA
Textbooks are all you needarxiv 2023-
Textbooks are all you need II: phi-1.5 technical reportarxiv 2023-
Is Child-Directed Speech Effective Training Data for Language Modelsarxiv 2024https://babylm.github.io/index.html
SciLitLLM: How to Adapt LLMs for Scientific Literature Understandingarxiv 2024https://github.com/dptech-corp/Uni-SMART
Anygpt: Unified multimodal llm with discrete sequence modelingarxiv 2024https://junzhan2000.github.io/AnyGPT.github.io/
Is synthetic data from generative models ready for image recognitionarxiv 2023https://github.com/CVMI-Lab/SyntheticData
Rephrasing the web: A recipe for compute and data-efficient language modelingarxiv 2024-
Physics of language models: Part 3.1, knowledge storage and extractionarxiv 2024https://physics.allen-zhu.com/part-3-knowledge/part-3-1
Llemma: An open language model for mathematicsarxiv 2023https://github.com/EleutherAI/math-lm
Enhancing multilingual language model with massive multilingual knowledge triplesarxiv 2021https://github.com/ntunlp/kmlm.git
Large language models, physics-based modeling, experimental measurements: the trinity of data-scarce learning of polymer propertiesarxiv 2024-

Fine-Tuning

PaperPublished inCode/Project
Self-Instruct: Aligning language models with self-generated instructionsarxiv 2023https://github.com/yizhongw/self-instruct
WizardLM: Empowering large language models to follow complex instructionsarxiv 2023https://github.com/nlpxucan/WizardLM
Code Llama: Open foundation models for codearxiv 2023https://github.com/meta-llama/codellama
Scaling Relationship on Learning Mathematical Reasoning with Large Language Modelsarxiv 2023https://github.com/OFA-Sys/gsm8k-ScRel
Self-Translate-Train: A Simple but Strong Baseline for Cross-lingual Transfer of Large Language Modelsarxiv 2024-
CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement LearningNeurIPS 2022https://github.com/salesforce/CodeRL
Self-play fine-tuning converts weak language models to strong language modelsarxiv 2024https://github.com/uclaml/SPIN
Language models can teach themselves to program betterarxiv 2022https://github.com/microsoft/PythonProgrammingPuzzles
DeepSeek-Prover: Advancing theorem proving in LLMs through large-scale synthetic dataarxiv 2024-
STaR: Bootstrapping reasoning with reasoningarxiv 2022-
Reinforced Self-Training (ReST) for Language Modelingarxiv 2023-
Beyond human data: Scaling self-training for problem-solving with language modelsarxiv 2023-
Code alpaca: An instruction-following llama model for code generationgithub 2023https://github.com/sahil280114/codealpaca
Stanford Alpaca: An Instruction-following LLaMA Modelgithub 2023https://github.com/tatsu-lab/stanford_alpaca
Huatuo: Tuning llama model with chinese medical knowledgearxiv 2023https://github.com/SCIR-HI/Huatuo-Llama-Med-Chinese
Magicoder: Source code is all you needarxiv 2023https://github.com/ise-uiuc/magicoder
Knowledge-Infused Prompting: Assessing and Advancing Clinical Text Data Generation with Large Language ModelsSTEPhttps://github.com/ritaranx/ClinGen
Unnatural instructions: Tuning language models with (almost) no human laborarxiv 2022https://github.com/orhonovich/unnatural-instructions
Baize: An open-source chat model with parameter-efficient tuning on self-chat dataarxiv 2023https://github.com/project-baize/baize-chatbot
Impossible Distillation for Paraphrasing and Summarization: How to Make High-quality Lemonade out of Small, Low-quality Modelarxiv 2023-
Llm2llm: Boosting llms with novel iterative data enhancementarxiv 2024https://github.com/SqueezeAILab/LLM2LLM
WizardCode: Empowering code large language models with Evol-Instructarxiv 2023https://github.com/nlpxucan/WizardLM
Generative AI for Math: Abelarxiv 2024-
Orca: Progressive learning from complex explanation traces of gpt-4arxiv 2023https://www.microsoft.com/en-us/research/project/orca/
Orca 2: Teaching small language models how to reasonarxiv 2023-
Mammoth: Building math generalist models through hybrid instruction tuningarxiv 2023https://tiger-ai-lab.github.io/MAmmoTH/
Lab: Large-scale alignment for chatbotsarxiv 2024-
Synthetic data (almost) from scratch: Generalized instruction tuning for language modelsarxiv 2024https://thegenerality.com/agi/
SciLitLLM: How to Adapt LLMs for Scientific Literature Understandingarxiv 2024https://github.com/dptech-corp/Uni-SMART/tree/main/SciLitLLM
Llava-med: Training a large language-and-vision assistant for biomedicine in one dayarxiv 2024https://github.com/microsoft/LLaVA-Med
Visual instruction tuningNIPS 2024-
Chartllama: A multimodal llm for chart understanding and generationarxiv 2023https://tingxueronghua.github.io/ChartLlama/
Sharegpt4v: Improving large multi-modal models with better captionsarxiv 2023https://sharegpt4v.github.io/
Next-gpt: Any-to-any multimodal llmarxiv 2023https://next-gpt.github.io/
Does synthetic data generation of llms help clinical text mining? arxiv 2023-
Ultramedical: Building specialized generalists in biomedicinearxiv 2024https://github.com/TsinghuaC3I/UltraMedical
Q: How to Specialize Large Vision-Language Models to Data-Scarce VQA Tasks? A: Self-Train on Unlabeled Images!arxiv 2023https://github.com/codezakh/SelTDA
MetaMeth: Bootstap your own mathematical questions for large language modelsarxiv 2024https://meta-math.github.io/
Symbol tuning improves in-context learning in language modelsarxiv 2023-
DISC-MedLLM: Bridging General Large Language Models and Real-World Medical Consultationarxiv 2023https://github.com/FudanDISC/DISC-MedLLM
Mathgenie: Generating synthetic data with question back-translation for enhancing mathematical reasoning of llmsarxiv 2024-
BianQue: Balancing the Questioning and Suggestion Ability of Health LLMs with Multi-turn Health Conversations Polished by ChatGPTarxiv 2023https://github.com/scutcyr/BianQue

Instruction-Tuning

PaperPublished inCode/Project
Alpaca: A Strong, Replicable Instruction-Following Modelhttps://github.com/tatsu-lab/stanford_alpaca
AlpaGasus: Training A Better Alpaca with Fewer DataarXiv 2023https://lichang-chen.github.io/AlpaGasus/
Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Qualityhttps://github.com/lm-sys/FastChat
WizardLM: Empowering Large Language Models to Follow Complex InstructionsarXiv 2023https://github.com/nlpxucan/WizardLM
Orca: Progressive Learning from Complex Explanation Traces of GPT-4arXiv 2023https://www.microsoft.com/en-us/research/project/orca/
Orca 2: Teaching Small Language Models How to ReasonarXiv 2023https://www.microsoft.com/en-us/research/project/orca/
Baize: An Open-Source Chat Model with Parameter-Efficient Tuning on Self-Chat DataarXiv 2023https://github.com/project-baize/baize-chatbot
LongForm: Effective Instruction Tuning with Reverse InstructionsarXiv 2023https://github.com/akoksal/LongForm
Visual Instruction TuningNeurIPS 2024https://llava-vl.github.io/
Improved Baselines with Visual Instruction TuningIEEE 2024https://llava-vl.github.io/
LLaVA-Plus: Learning to Use Tools for Creating Multimodal AgentsarXiv 2023https://llava-vl.github.io/llava-plus/
LLaVA-Interactive: An All-in-One Demo for Image Chat, Segmentation, Generation and EditingarXiv 2023-
LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One DayNeurIPS 2024https://aka.ms/llava-med
Self-Instruct: Aligning Language Models with Self-Generated InstructionsarXiv 2022https://github.com/yizhongw/self-instruct
Self-Alignment with Instruction BacktranslationarXiv 2023
Self-Play Fine-Tuning Converts Weak Language Models to Strong Language ModelsarXiv 2024https://github.com/uclaml/SPIN
Beyond Human Data: Scaling Self-Training for Problem-Solving with Language ModelsarXiv 2023-
Constitutional AI: Harmlessness from AI FeedbackarXiv 2022-
Toolformer: Language Models Can Teach Themselves to Use ToolsarXiv 2023-
Q: How to Specialize Large Vision-Language Models to Data-Scarce VQA Tasks? A: Self-Train on Unlabeled Images!CVPR 2023https://github.com/codezakh/SelTDA
ChatGPT-4 Outperforms Experts and Crowd Workers in Annotating Political Twitter Messages with Zero-Shot LearningarXiv 2023-
Prompting Large Language Model for Machine Translation: A Case StudyarXiv 2023-
T-SciQ: Teaching Multimodal Chain-of-Thought Reasoning via Mixed Large Language Model Signals for Science Question AnsweringAAAI 2024https://github.com/T-SciQ/T-SciQ
CORE: A Retrieve-then-Edit Framework for Counterfactual Data GenerationarXiv 2022https://github.com/tanay2001/CORE
Diversify Your Vision Datasets with Automatic Diffusion-Based AugmentationNeurIPS 2023https://github.com/lisadunlap/ALIA
AugGPT: Leveraging ChatGPT for Text Data AugmentationarXiv 2023-
CoAnnotating: Uncertainty-Guided Work Allocation between Human and Large Language Models for Data AnnotationarXiv 2023https://github.com/SALT-NLP/CoAnnotating
Closing the Loop: Testing ChatGPT to Generate Model Explanations to Improve Human Labelling of Sponsored Content on Social MediaSpringer, Cham-
ToolCoder: Teach Code Generation Models to use API search toolsarXiv 2023-

Preference Alignment

PaperPublished inCode/Project
UltraFeedback: Boosting Language Models with Scaled AI FeedbackarXiv 2023-
HelpSteer: Multi-attribute Helpfulness Dataset for SteerLMarXiv 2023https://huggingface.co/datasets/nvidia/HelpSteer
Learning From Mistakes Makes LLM Better ReasonerarXiv 2023https://github.com/microsoft/LEMA
Bot-Adversarial Dialogue for Safe Conversational AgentsACL 2021https://parl.ai/projects/safety_recipes/
BeaverTails: Towards Improved Safety Alignment of LLM via a Human-Preference DatasetNIPS 2024https://sites.google.com/view/pku-beavertails
Let's Verify Step by SteparXiv 2023-
WebGPT: Browser-assisted question-answering with human feedbackarXiv 2021https://www.microsoft.com/en-us/bing/apis/bing-web-search-api
Direct Language Model Alignment from Online AI FeedbackarXiv 2024-
Self-Judge: Selective Instruction Following with Alignment Self-EvaluationarXiv 2024https://github.com/nusnlp/Self-J
SALMON: Self-Alignment with Instructable Reward ModelsICLR 2024https://github.com/IBM/SALMON
SteerLM: Attribute Conditioned SFT as an (User-Steerable) Alternative to RLHFarXiv 2023https://huggingface.co/nvidia/SteerLM-llama2-13B
Starling-7B: Increasing LLM Helpfulness & Harmlessness with RLAIFCOLM 2024https://starling.cs.berkeley.edu/
Advancing LLM Reasoning Generalists with Preference TreesarXiv 2024https://github.com/OpenBMB/Eurus
CriticBench: Benchmarking LLMs for Critique-Correct ReasoningarXiv 2024https://criticbench.github.io/

Applications

Math

PaperPublished inCode/Project
Galactica: A Large Language Model for Sciencearxiv 2022-
STaR: Bootstrapping Reasoning With ReasoningNeurIPS 2022https://github.com/ezelikman/STaR
Multilingual Mathematical Autoformalizationarxiv 2023-
WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instructarxiv 2023https://github.com/nlpxucan/WizardLM
MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuningarxiv 2023-
MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Modelsarxiv 2023https://meta-math.github.io/
Synthetic Dialogue Dataset Generation using LLM AgentsEMNLP Workshop 2023-
Advancing Theorem Proving in LLMs through Large-Scale Synthetic DataNeurIPS Workshop 2024-
Synthetic Dialogue Dataset Generation using LLM Agentsarxiv 2024-
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Modelsarxiv 2024https://github.com/deepseek-ai/DeepSeek-Math
DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Dataarxiv 2024-
Augmenting Math Word Problems via Iterative Question Composingarxiv 2024https://huggingface.co/datasets/Vivacem/MMIQC

Science

PaperPublished inCode/Project
Galactica: A Large Language Model for Sciencearxiv 2022-
Reflection-Tuning: Recycling Data for Better Instruction-TuningNeurIPS Workshop 2023 / ACL 2024https://github.com/tianyi-lab/Reflection_Tuning
Reflexion: language agents with verbal reinforcement learningNeurIPS 2023https://github.com/noahshinn024/reflexion
SciLitLLM: How to Adapt LLMs for Scientific Literature UnderstandingNeurIPS Workshop 2024https://github.com/dptech-corp/Uni-SMART/tree/main/SciLitLLM
SciInstruct: a Self-Reflective Instruction Annotated Dataset for Training Scientific Language ModelsNeurIPS 2024https://github.com/THUDM/SciGLM
ChemLLM: A Chemical Large Language Modelarxiv 2024-

Code

PaperPublished inCode/Project
CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement LearningNIPS 2022https://github.com/salesforce/CodeRL
Generating Programming Puzzles to Train Language ModelsICLR 2022 (Workshop)-
Language Models Can Teach Themselves to Program BetterICLR 2023-
Textbooks Are All You NeedArxiv 2023-
Textbooks Are All You Need II: phi-1.5 technical reportArxiv 2023-
Self-Consistency Improves Chain of Thought Reasoning in Language ModelsICLR 2023-
Learning Performance-Improving Code EditsICLR 2024https://pie4perf.com/
WizardCoder: Empowering Code Large Language Models with Evol-InstructICLR 2024https://github.com/nlpxucan/WizardLM
Magicoder: Source Code Is All You NeedICML 2024https://github.com/ise-uiuc/magicoder

Medical

PaperPublished inCode/Project
MedDialog: Large-scale Medical Dialogue DatasetsEMNLP 2020https://github.com/UCSDAI4H/Medical-Dialogue-System
HuatuoGPT, towards Taming Language Model to Be a DoctorEMNLP 2023https://github.com/FreedomIntelligence/HuatuoGPT
HuatuoGPT-II, One-stage Training for Medical Adaption of LLMsarxiv 2023https://github.com/FreedomIntelligence/HuatuoGPT-II
ChatCounselor: A Large Language Models for Mental Health Supportarxiv 2023-
DISC-MedLLM: Bridging General Large Language Models and Real-World Medical Consultationarxiv 2023https://github.com/FudanDISC/DISC-MedLLM
Biomedical discovery through the integrative biomedical knowledge hub (iBKH)iScience 2023https://github.com/wcm-wanglab/iBKH
Knowledge-Infused Prompting: Assessing and Advancing Clinical Text Data Generation with Large Language Modelsarxiv 2023https://github.com/ritaranx/ClinGen
ShenNong-TCMGithub repohttps://github.com/michael-wzhu/ShenNong-TCM-LLM
ZhongJing(仲景)Github repohttps://github.com/pariskang/CMLM-ZhongJing

Law

PaperPublished inCode/Project
DISC-LawLLM: Fine-tuning Large Language Models for Intelligent Legal Servicesarxiv 2023https://github.com/FudanDISC/DISC-LawLLM
Lawyer LLaMA Technical Reportarxiv 2023-
LawGPT: A Chinese Legal Knowledge-Enhanced Large Language Modelarxiv 2024https://github.com/pengxiao-song/LaWGPT
WisdomInterrogatoryGithub repohttps://github.com/zhihaiLLM/wisdomInterrogatory

Education

PaperPublished inCode/Project
A Comparative Study of AI-Generated (GPT-4) and Human-crafted MCQs in Programming EducationProceedings of the 26th Australasian Computing Education Conference-

Financial

PaperPublished inCode/Project
FinTral: A Family of GPT-4 Level Multimodal Financial Large Language ModelsArxiv 2024http://arxiv.org/abs/2402.10986
FinGLM CompetitionGithub repohttps://github.com/MetaGLM/FinGLM

Functionality

Understanding

PaperPublished inCode/Project
Alpaca: A Strong, Replicable Instruction-Following Model-https://github.com/tatsu-lab/stanford_alpaca
WizardLM: Empowering Large Language Models to Follow Complex Instructionsarxiv 2023https://github.com/nlpxucan/WizardLM
Rephrasing the Web: A Recipe for Compute and Data-Efficient Language Modelingarxiv 2024-
Visual Instruction TuningNIPS 2024https://llava-vl.github.io
ChartLlama: A Multimodal LLM for Chart Understanding and Generationarxiv 2023https://tingxueronghua.github.io/ChartLlama/
Genixer: Empowering Multimodal Large Language Models as a Powerful Data Generatorarxiv 2023https://github.com/zhaohengyuan1/Genixer

Logic

PaperPublished inCode/Project
Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Modelsarxiv 2023-
Case2Code: Learning Inductive Reasoning with Synthetic Dataarxiv 2024https://github.com/choosewhatulike/case2code
MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuningarxiv 2023https://tiger-ai-lab.github.io/MAmmoTH/
Augmenting Math Word Problems via Iterative Question ComposingICLR 2024https://huggingface.co/datasets/Vivacem/MMIQC
STaR: Bootstrapping Reasoning With Reasoningarxiv 2022-
Q: How to Specialize Large Vision-Language Models to Data-Scarce VQA Tasks? A: Self-Train on Unlabeled Images!CVPR 2023https://github.com/codezakh/SelTDA

Memory

PaperPublished inCode/Project
Quiet-STaR: Language Models Can Teach Themselves to Think Before Speakingarxiv 2024-
LLMs for Knowledge Graph Construction and Reasoning: Recent Capabilities and Future OpportunitiesWorld Wide Webhttps://github.com/zjunlp/AutoKG
Scaling Synthetic Data Creation with 1,000,000,000 Personasarxiv 2024https://github.com/tencent-ailab/persona-hub
AceCoder: Utilizing Existing Code to Enhance Code Generationarxiv 2023-
RepoCoder: Repository-Level Code Completion Through Iterative Retrieval and GenerationarXiv 2023https://github.com/microsoft/CodeT/tree/main/RepoCoder

Generation

PaperPublished inCode/Project
Genie: Achieving Human Parity in Content-Grounded Datasets Generationarxiv 2024-
UltraMedical: Building Specialized Generalists in Biomedicinearxiv 2024https://github.com/TsinghuaC3I/UltraMedical
HuaTuo: Tuning LLaMA Model with Chinese Medical Knowledgearxiv 2023https://github.com/SCIR-HI/Huatuo-Llama-Med-Chinese
TinyStories: How Small Can Language Models Be and Still Speak Coherent English?arxiv 2023-
Controllable Dialogue Simulation with In-Context Learningarxiv 2022https://github.com/Leezekun/dialogic
Diversify Your Vision Datasets with Automatic Diffusion-Based AugmentationNIPS 2023https://github.com/lisadunlap/ALIA

Challenges and Limitations

Synthesizing and Augmenting Method

PaperPublished inCode/Project
RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedbackarxiv 2023-
LLM2LLM: Boosting LLMs with Novel Iterative Data Enhancementarxiv 2024https://github.com/SqueezeAILab/LLM2LLM
WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instructarxiv 2023https://github.com/nlpxucan/WizardLM
STaR: Bootstrapping Reasoning With ReasoningNIPS 2022-
SciInstruct: a Self-Reflective Instruction Annotated Dataset for Training Scientific Language Modelsarxiv 2024https://github.com/THUDM/SciGLM
ChemLLM: A Chemical Large Language Modelarxiv 2024-

Data Quality

PaperPublished inCode/Project
LLMs4Synthesis: Leveraging Large Language Models for Scientific Synthesisarxiv 2024-
CoRAL: Collaborative Retrieval-Augmented Large Language Models Improve Long-tail RecommendationACM-
Examining Inter-Consistency of Large Language Models Collaboration: An In-depth Analysis via Debatearxiv 2023-
LTGC: Long-tail Recognition via Leveraging LLMs-driven Generated ContentCVPR 2024https://ltgccode.github.io

Impact of Data Synthesis and Augmentation

PaperPublished inCode/Project
DataDreamer: A Tool for Synthetic Data Generation and Reproducible LLM Workflowsarxiv 2024https://github.com/datadreamer-dev/DataDreamer
HARMONIC: Harnessing LLMs for Tabular Data Synthesis and Privacy Protectionarxiv 2024https://github.com/The-FinAI/HARMONIC

Impact on Different Applications and Tasks

PaperPublished inCode/Project
PANDA: Preference Adaptation for Enhancing Domain-Specific Abilities of LLMsACL 2024https://github.com/THUNLP-MT/PANDA
Role Prompting Guided Domain Adaptation with General Capability Preserve for Large Language Modelsarxiv 2024-

Future Directions

PaperPublished inCode/Project
A Universal Metric for Robust Evaluation of Synthetic Tabular DataIEEE 2022-
CoLa-Diff: Conditional Latent Diffusion Model for Multi-Modal MRI SynthesisSpringer 2023-
LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agentsarxiv 2023https://llava-vl.github.io/llava-plus/
WizardCoder: Empowering Code Large Language Models with Evol-Instructarxiv 2023https://github.com/nlpxucan/WizardLM
AnyGPT: Unified Multimodal LLM with Discrete Sequence Modelingarxiv 2024https://junzhan2000.github.io/AnyGPT.github.io/
WebGPT: Browser-assisted question-answering with human feedbackarxiv 2021https://www.microsoft.com/en-us/bing/apis/bing-web-search-api
NExT-GPT: Any-to-Any Multimodal LLMarxiv 2023https://next-gpt.github.io/