DeepSeek-V4-Flash-Vision-Exp deepseek-ai
- 種別
- model
- ライセンス
- mit
- タスク
- image-text-to-text
- フレームワーク
- transformers
- ダウンロード
- 0
- いいね
- 259
- パラメータ
- 304,646,824,126
- アクセス
- public
- ファイル
- 84
タグ
- 多模态
- 大言語モデル
- 視覚言語モデル
- エージェント
- 画像理解
- MoE
- マルチモーダル
- PyTorch
概要
DeepSeek-V4シリーズ初の実験的多モーダルモデルで、DeepSeek-V4-Flashにビジョンモジュールを統合し継続学習することで画像理解とマルチモーダルエージェント能力を獲得した。テキストのみのエージェント性能を維持しつつ、ApexBench等のマルチモーダルエージェントベンチマークで大幅な改善を示す。vision encoder・MoE・Hyper-Connectionsを備えたPyTorch推論実装とプロンプトエンコーディングを提供し、Terminal BenchやChartographyなどのエージェントタスクに適用可能。
README
--- license: mit library_name: transformers pipeline_tag: image-text-to-text --- # DeepSeek-V4-Flash-Vision-Exp <!-- markdownlint-disable first-line-h1 --> <!-- markdownlint-disable html --> <!-- markdownlint-disable no-duplicate-header --> <div align="center"> <img src="https://github.com/deepseek…