Learning Transferable Models for Computer Vision via Image–Text Pre-training¶
Radford, A., Kim, & Hallacy. (2021). Learning Transferable Models for Computer Vision via Image–Text Pre-training. Proceedings of the 38th International Conference on Machine Learning.
Cited by¶
2 citations across 2 artifacts.
Each citation links to the sentence it supports in the citing article.
Primes¶
- Icon–Index–Symbol Distinction
- T6b — Multimodal AI and Peircean sign-types. Recent multimodal models (CLIP, LLaVA, GPT-4V) implement iconic (image-matching), indexical (correlation-detection), and symbolic (language-based) operations simultaneously.
This sourceRadford CLIP multimodal vision-language iconic indexical symbolic AI.
- T6b — Multimodal AI and Peircean sign-types. Recent multimodal models (CLIP, LLaVA, GPT-4V) implement iconic (image-matching), indexical (correlation-detection), and symbolic (language-based) operations simultaneously.
- Iconicity
- Multimodal AI and vision-language models: Radford et al. (2021, CLIP)
This sourceRadford CLIP multimodal vision-language model implicit iconic alignment.
- Multimodal AI and vision-language models: Radford et al. (2021, CLIP)
Verification¶
This reference passed the adversarial substantiation pipeline: it was checked to exist and to support the claim it is attached to. See how references were verified.
Registry ID ref:33b9e2f4abd8 · see in the full table