The large model of "Scholar Puyu Ling Pen" in Shanghai artificial Intelligence Laboratory is officially open source.
CTOnews.com, October 10 (Xinhua)-- Shanghai artificial Intelligence Lab has launched its first large picture-text hybrid model, Puyu Ling Pen (InternLM-XComposer), and announced that it is open source, and has launched GitHub, Hugging Face and Munda Community at the same time.
According to reports, the Pu language Ling Pen is based on the Pu language language Model (InternLM). It has a strong multimodal performance, can accept visual and language modal input, and can also generate mixed text articles with one click.
It is worth mentioning that the researchers used five mainstream multimodal large model evaluations to test the capabilities of InternLM-XComposer-VL-7B in detail, including:
MME Benchmark: comprehensive evaluation of the multimodal model including 14 subtasks, focusing on the model's Perception and Recognition capabilities
MMBench: including 20 latitudes of capability and multimodal evaluation using the ChatGPT cycle evaluation strategy
MMBench-CN: MMBench Evaluation of questions and answers in simplified Chinese version
Seed-Bench: provides multimodal evaluation of 1.9 million channels of multi-modal and multi-choice topics including manual labeling.
CCBench: a Chinese multimodal assessment of Chinese cultural understanding. The evaluation results show that the Pu language Ling pen shows excellent performance in the above five Chinese and English multimodal evaluations.
At present, Pu language Ling Pen has opened up intelligent creation and dialogue (InternLM-XComposer-7B) and multi-task pre-training (InternLM-XComposer-VL-7B) versions, and provides free commercial use. CTOnews.com with official address:
Open source link: https://github.com/ InternLM / InternLM-XComposer
Technical report: https://arxiv.org/ abs / 2309.15112