Upload Records Snowball Search Search OpenAlex
About the database ScholarIQanswers from OpenAlex
Multimodal Machine Learning Applications
TopicLeading institutions, researchers & key papers
This cluster of papers focuses on the development and improvement of visual question answering systems, image captioning techniques, and neural networks for understanding and generating descriptions of images and videos. The research involves semantic reasoning, multimodal fusion, scene graph generation, attention mechanisms, and deep learning approaches to bridge the gap between vision and language.
106
Works
IDs:OpenAlex
How has Multimodal Machine Learning Applications's publication output changed over time?
ScholarIQpublication output · 2015–2022
Output grew0% over the shown period — from 2 works in 2015 to 2 in 2022.
2
1
3
2
2
2
2
2015201620172019202020212022
What are the most-cited papers on Multimodal Machine Learning Applications?
ScholarIQmost cited works
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, Li Fei-Fei
S25538012. 20175,247 CitationsOPEN ACCESS
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy, Li Fei-Fei
20155,006 Citations
A Comprehensive Survey of Deep Learning for Image Captioning
Md Zakir Hossain, Ferdous Sohel, Mohd Fairuz Shiratuddin, Hamid Laga
S157921468. 2019887 Citations
VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language Research
Xin Wang, Jiawei Wu, Junkun Chen, Lei Li, Yuan‐Fang Wang, William Yang Wang
2019474 Citations
GLoRIA: A Multimodal Global-Local Representation Learning Framework for Label-efficient Medical Image Recognition
Shih-Cheng Huang, Liyue Shen, Matthew P. Lungren, Serena Yeung
S4363607764. 2021407 Citations
Where is Multimodal Machine Learning Applications research published, and who funds it?
ScholarIQvenues & funding sources
TOP JOURNALS
TOP FUNDERS
National Science Foundation—
NIH—
Wellcome Trust—
European Research Council—
Funder breakdown is a member featureSign up free to unlock
How much of the research on Multimodal Machine Learning Applications is open access?
ScholarIQopen access share
27%OPEN ACCESS
Gold
7%
Green
13%
Hybrid
7%
Bronze
0%
Closed
73%
Related on ScholarIQ
ImageNet: A large-scale hierarchical image database
Paper
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
Paper
Deep visual-semantic alignments for generating image descriptions
Paper
TinyBERT: Distilling BERT for Natural Language Understanding
Paper
A Comprehensive Survey of Deep Learning for Image Captioning
Paper
End-to-End Learning of Action Detection from Frame Glimpses in Videos
Paper