Get the App
SLTechnology News&Howtos  ›  IT Information  › 

Beijing Zhiyuan launched a general visual AI model SegGPT: it can automatically track and segment objects in audio and video.

Shulou Source: shulou.com Published: 2023-11-24 17:09:49 10月02日 Update

Thanks to CTOnews.com netizen Xiao Zhan cut the clue delivery! CTOnews.com May 31 news, in 2023 Zhongguancun Forum artificial intelligence large model development forum, Beijing Zhiyuan artificial intelligence research institute launched its general segmentation model Segment Everything In Contex (GPT).

▲ Image source Arxiv It is said that SegGPT model is a derivative model of Painter, which has contextual reasoning ability. After training, it only needs to provide examples to reason and complete corresponding segmentation tasks, including examples, categories, parts, contours, texts, faces, medical images, etc. in images and videos, all of which can be segmented by visual prompts.

Arxiv SegGPT also has reasoning capabilities that support any number of visual cues. Automatic video segmentation can be performed with the first frame image and corresponding object mask as context examples, and automatic tracking can be performed with the color of the mask as the ID of the object.

CTOnews.com learned from a query that Meta has also released its AI-based Segment Anything Model (SAM), which has the ability to identify and separate specific objects in images and videos. Researchers from Wisconsin Madison, Microsoft, Hong Kong University of Science and Technology have also launched SEEM models to segment images and videos with one click through different visual cues and language cues. CTOnews.com friends can access the model's paper link here.

Tags: Models images vision video prompts capabilities reasoning objects up and down context artificial intelligence tasks intelligence examples forums research Beijing different people Apple Docker Huawei Linux macOS MariaDB Microsoft MySQL NVidia OPPO Reno Linux MySQL Microsoft OPPO Reno Redmi