Scholar IQ
Try ScholarIQ free
Upload Records Snowball Search Search OpenAlex
About the database
On this page:OverviewPublicationsResearchersKey papersJournalsOpen accessInstitutions
ScholarIQanswers from OpenAlex

Multimodal Machine Learning Applications

TopicLeading institutions, researchers & key papers

This cluster of papers focuses on the development and improvement of visual question answering systems, image captioning techniques, and neural networks for understanding and generating descriptions of images and videos. The research involves semantic reasoning, multimodal fusion, scene graph generation, attention mechanisms, and deep learning approaches to bridge the gap between vision and language.

106
Works

How has Multimodal Machine Learning Applications's publication output changed over time?

ScholarIQpublication output · 2015–2022

Output grew0% over the shown period — from 2 works in 2015 to 2 in 2022.

2
1
3
2
2
2
2
2015201620172019202020212022

What are the most-cited papers on Multimodal Machine Learning Applications?

ScholarIQmost cited works
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, Li Fei-Fei
S25538012. 20175,247 CitationsOPEN ACCESS
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy, Li Fei-Fei
20155,006 Citations
A Comprehensive Survey of Deep Learning for Image Captioning
Md Zakir Hossain, Ferdous Sohel, Mohd Fairuz Shiratuddin, Hamid Laga
S157921468. 2019887 Citations
VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language Research
Xin Wang, Jiawei Wu, Junkun Chen, Lei Li, Yuan‐Fang Wang, William Yang Wang
2019474 Citations
GLoRIA: A Multimodal Global-Local Representation Learning Framework for Label-efficient Medical Image Recognition
Shih-Cheng Huang, Liyue Shen, Matthew P. Lungren, Serena Yeung
S4363607764. 2021407 Citations

Where is Multimodal Machine Learning Applications research published, and who funds it?

ScholarIQvenues & funding sources

TOP JOURNALS

S255380125,247
S157921468887
S4363607764407
S7560371206

TOP FUNDERS

National Science Foundation
NIH
Wellcome Trust
European Research Council
Funder breakdown is a member featureSign up free to unlock

How much of the research on Multimodal Machine Learning Applications is open access?

ScholarIQopen access share
27%OPEN ACCESS
Gold
7%
Green
13%
Hybrid
7%
Bronze
0%
Closed
73%

Related on ScholarIQ

ImageNet: A large-scale hierarchical image database
Paper
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
Paper
Deep visual-semantic alignments for generating image descriptions
Paper
TinyBERT: Distilling BERT for Natural Language Understanding
Paper
A Comprehensive Survey of Deep Learning for Image Captioning
Paper
End-to-End Learning of Action Detection from Frame Glimpses in Videos
Paper
470M+ articles · free account