Yunnglin/Yunnglin
Config files for my GitHub profile.
Master's graduate from Nanjing University (@NJUNLP), currently working at @ModelScope.
Config files for my GitHub profile.
[ICML 2025] MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding
This repo contains evaluation code for the paper "MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI"
MathVista: data, code, and evaluation for Mathematical Reasoning in Visual Contexts
[CVPR 2025] A Comprehensive Benchmark for Document Parsing and Evaluation
Repository for getting started with the OfficeQA Benchmark.
[ICLR 2026] The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution
Claw-Eval is an evaluation harness for evaluating LLM as agents. All tasks verified by humans.
WideSearch: Benchmarking Agentic Broad Info-Seeking
MCore-Bridge: Providing Megatron-Core model definitions for state-of-the-art large models and making Megatron training as simple as Transformers — with support for 300+ large language models (Qwen3-Next, GLM-5.1, Deepseek-V4, MiniMax-2.7, ...) and 200+ multimodal large models (Qwen3.5, Qwen3-Omni, Gemma4, ...).
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
Modularized and Stable Sandbox runtime environment
Enjoy the magic of Diffusion models!
🐫 CAMEL: The first and the best multi-agent framework. Finding the Scaling Law of Agents. https://www.camel-ai.org
Langflow is a low-code app builder for RAG and multi-agent AI applications. It’s Python-based and agnostic to any model, API, or database.
Open-source evaluation toolkit of large vision-language models (LVLMs), support GPT-4v, Gemini, QwenVLPlus, 50+ HF models, 20+ benchmarks
🦜🔗 Build context-aware reasoning applications
Qwen2.5 is the large language model series developed by Qwen team, Alibaba Cloud.
ModelScope: bring the notion of Model-as-a-Service to life.
Supercharge Your LLM Application Evaluations 🚀
OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets.
ms-swift: Use PEFT or Full-parameter to finetune 300+ LLMs or 50+ MLLMs. (Qwen2, GLM4v, Internlm2.5, Yi, Llama3, Llava-Video, Internvl2, MiniCPM-V, Deepseek, Baichuan2, Gemma2, Phi3-Vision, ...)
demo for learning github actions
🤗 Transformers: State-of-the-art Machine Learning for Pytorch, TensorFlow, and JAX.
python server for go-cqhttp