analytics
data-juicer: Simplify Data Processing
Founders get easy data processing for foundation models with data-juicer, a Python tool with 6.8k+ GitHub stars, ideal for data-driven startups.
6,794 stars400 forksPythonHealth Score 9/10Updated 7/28/2026100% free · open source
What it does
Data-juicer is a tool for processing data for and with foundation models, allowing users to efficiently handle and transform data for machine learning tasks
Install / run
pip install data-juicerWhen to use it
- •When you need to preprocess large datasets for training foundation models
- •When you want to integrate data processing with existing foundation model pipelines
- •When you require flexible and customizable data transformation for specific machine learning tasks
Quick start
- 1 Clone the data-juicer repository from GitHub using the command `git clone https://github.com/datajuicer/data-juicer`
- 2Navigate to the cloned repository using `cd data-juicer`
- 3Install the required packages using `pip install -r requirements.txt`
- 4Run the example script using `python examples/example.py` to test the installation
- 5Modify the `config.json` file to customize the data processing pipeline for your specific use case
Ready-to-paste prompt
python data_juicer.py --input_path data/input.csv --output_path data/output.csv --model_name 'my_foundation_model'
Heads up: Make sure you have the correct version of Python installed, as data-juicer is compatible with Python 3.8 and later, and also ensure you have the necessary dependencies installed, including pandas and numpy
Saves to your device
Topics
data
data-analysis
data-pipeline
data-processing
data-science
data-visualization
foundation-models
instruction-tuning
large-language-models
llm
llms
multi-modal
pre-training
synthetic-data
What's inside — free to inspect
No purchase needed
Read the entire source before you build — unlike paid marketplaces that hide it behind a buy button.
15
top-level files
9
folders
1972.2M
source size
Apache-2.0
license
Key files
.pre-commit-config.yaml
README_ZH.md
README.md
File tree
.github/
.pre-commit-hooks/
data_juicer/
demos/
docs/
scripts/
tests/
thirdparty/
tools/
.coveragerc
.gitignore
.pre-commit-config.yaml
.secrets.baseline
app.py
Dockerfile
Dockerfile.embodied
hatch_build.py
label_studio_localhost_connection.json
LICENSE
pyproject.toml
README_ZH.md
README.md
service.py
uv.lock
Details
Creator
datajuicer
Language
Python
Category
analytics
Published
8/1/2023
Are you the creator of this tool? Claim your listing → and earn 85% of every sale.
Related skills
More analytics tools founders pair with this one.
analytics★ 110,200
TypeScript: Cleaner JS Code
Get reliable JavaScript output with TypeScript. For founders using JavaScript.
analytics★ 90,206
Vue Element Admin
:tada: A magical vue admin https://panjiachen.github.io/vue-element-admin
analytics★ 76,261
Grafana: Unified Insights
Get unified metrics and monitoring for your startup with Grafana, a platform used by many, ideal for founders needing data visibility.
analytics★ 74,301
Superset: Data Insights
Get data visualization and exploration with Superset, a platform for founders in data-driven startups, with 74k+ GitHub stars.
analytics★ 62,997
Daily Stock Analysis
Get intelligent A/H/US market insights with daily_stock_analysis. For founders needing AI-driven stock analysis.
analytics★ 48,809
Metabase: Easy Data Insights
Get data-driven decisions with Metabase, an open source BI tool for founders, with 48k+ GitHub stars.