Yunlin Mao

@Yunnglin · User

GitHub profile ↗ · Compare

Master's graduate from Nanjing University (@NJUNLP), currently working at @ModelScope.

Tongyi Lab, Alibaba Group60 followers55 repositories

Repositories

Yunnglin/MedXpertQA

[ICML 2025] MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

★ 0Forks 0

Yunnglin/Video-MME-v2

Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding

★ 0Forks 0

Yunnglin/MMMU

This repo contains evaluation code for the paper "MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI"

★ 0Forks 0

Yunnglin/MathVista

MathVista: data, code, and evaluation for Mathematical Reasoning in Visual Contexts

★ 0Forks 0

Yunnglin/OmniDocBench

[CVPR 2025] A Comprehensive Benchmark for Document Parsing and Evaluation

★ 0Forks 0

Yunnglin/officeqa

Repository for getting started with the OfficeQA Benchmark.

★ 0Forks 0

Yunnglin/Toolathlon

[ICLR 2026] The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution

★ 0Forks 0

Yunnglin/claw-eval

Claw-Eval is an evaluation harness for evaluating LLM as agents. All tasks verified by humans.

★ 0Forks 0

Yunnglin/mcore-bridge

MCore-Bridge: Providing Megatron-Core model definitions for state-of-the-art large models and making Megatron training as simple as Transformers — with support for 300+ large language models (Qwen3-Next, GLM-5.1, Deepseek-V4, MiniMax-2.7, ...) and 200+ multimodal large models (Qwen3.5, Qwen3-Omni, Gemma4, ...).

★ 0Forks 0

Yunnglin/ray

Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.

★ 0Forks 0

Yunnglin/camel

🐫 CAMEL: The first and the best multi-agent framework. Finding the Scaling Law of Agents. https://www.camel-ai.org

★ 0PythonForks 0

Yunnglin/langflow

Langflow is a low-code app builder for RAG and multi-agent AI applications. It’s Python-based and agnostic to any model, API, or database.

★ 0PythonForks 0

Yunnglin/VLMEvalKit

Open-source evaluation toolkit of large vision-language models (LVLMs), support GPT-4v, Gemini, QwenVLPlus, 50+ HF models, 20+ benchmarks

★ 0PythonForks 0

Yunnglin/Qwen2.5

Qwen2.5 is the large language model series developed by Qwen team, Alibaba Cloud.

★ 0ShellForks 0

Yunnglin/modelscope

ModelScope: bring the notion of Model-as-a-Service to life.

★ 0PythonForks 0

Yunnglin/ragas

Supercharge Your LLM Application Evaluations 🚀

★ 0PythonForks 0

Yunnglin/opencompass

OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets.

★ 0PythonForks 0

Yunnglin/swift

ms-swift: Use PEFT or Full-parameter to finetune 300+ LLMs or 50+ MLLMs. (Qwen2, GLM4v, Internlm2.5, Yi, Llama3, Llava-Video, Internvl2, MiniCPM-V, Deepseek, Baichuan2, Gemma2, Phi3-Vision, ...)

★ 0PythonForks 0

Yunnglin/transformers

🤗 Transformers: State-of-the-art Machine Learning for Pytorch, TensorFlow, and JAX.

★ 0Forks 0