DeepSeek-V4-Flash-Vision-Exp deepseek-ai

種別
model
ライセンス
mit
タスク
image-text-to-text
フレームワーク
transformers
ダウンロード
0
いいね
259
パラメータ
304,646,824,126
アクセス
public
ファイル
84

タグ

  • 多模态
  • 大言語モデル
  • 視覚言語モデル
  • エージェント
  • 画像理解
  • MoE
  • マルチモーダル
  • PyTorch

概要

DeepSeek-V4シリーズ初の実験的多モーダルモデルで、DeepSeek-V4-Flashにビジョンモジュールを統合し継続学習することで画像理解とマルチモーダルエージェント能力を獲得した。テキストのみのエージェント性能を維持しつつ、ApexBench等のマルチモーダルエージェントベンチマークで大幅な改善を示す。vision encoder・MoE・Hyper-Connectionsを備えたPyTorch推論実装とプロンプトエンコーディングを提供し、Terminal BenchやChartographyなどのエージェントタスクに適用可能。

README

--- license: mit library_name: transformers pipeline_tag: image-text-to-text --- # DeepSeek-V4-Flash-Vision-Exp <!-- markdownlint-disable first-line-h1 --> <!-- markdownlint-disable html --> <!-- markdownlint-disable no-duplicate-header --> <div align="center"> <img src="https://github.com/deepseek…

查看完整页面 · 查看原文