1. X
  2. Gabriel Ilharco
Log inSign up
Gabriel Ilharco
432 posts
Image
user avatar
Gabriel Ilharco
@gabriel_ilharco
AI Research Scientist
Palo Alto, CA
gabrielilharco.com
Joined September 2015
1,330
Following
6,495
Followers
RepliesRepliesMediaMedia

New to X?

Sign up now to get your own personalized timeline!

Create account

By signing up, you agree to the Terms of Service and Privacy Policy, including Cookie Use.

Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Don't miss what's happening
People on X are the first to know.
Log inSign up
  • user avatar
    Gabriel Ilharco
    @gabriel_ilharco
    Dec 15, 2022
    Introducing task vectors! A new way to steer models by doing arithmetic with model weights. Subtract to make models forget, add to make them learn 📜: arxiv.org/abs/2212.04089 🖥️: github.com/mlfoundations/…
    Image
  • user avatar
    Gabriel Ilharco
    @gabriel_ilharco
    Apr 28, 2023
    Introducing DataComp, a new benchmark for multimodal datasets! We release 12.8B image-text pairs, 300+ experiments and a 1.4B subset that outcompetes compute-matched CLIP runs from OpenAI & LAION 📜 arxiv.org/abs/2304.14108 🖥️ github.com/mlfoundations/… 🌐 datacomp.ai
    Image
    208K0208K
  • user avatar
    Gabriel Ilharco
    @gabriel_ilharco
    May 4, 2023
    Today we are releasing a CLIP ViT-L/14 model with 79.2% zero-shot accuracy on ImageNet. Our model outperforms OpenAI's CLIP by a large margin, and outperforms even bigger models (ViT-g/14) trained on LAION-2B Check it out at huggingface.co/laion/CLIP-ViT…!
    Image
    laion/CLIP-ViT-L-14-DataComp.XL-s13B-b90K · Hugging Face
    From huggingface.co
    190K0190K
  • user avatar
    Gabriel Ilharco
    @gabriel_ilharco
    Oct 26, 2023
    CLIP models have become a lot better since 2021
    Image
    161K0161K
  • user avatar
    Gabriel Ilharco
    @gabriel_ilharco
    Mar 3, 2022
    Fine-tuning can make models like CLIP less robust. A simple idea is highly effective at mitigating that: averaging zero-shot and fine-tuned models. Check out our work introducing WiSE-FT, just accepted to CVPR! Paper: arxiv.org/abs/2109.01903 Code: github.com/mlfoundations/…
    Image
  • user avatar
    Gabriel Ilharco
    @gabriel_ilharco
    Jul 29, 2021
    We are releasing an open-source training implementation of OpenAI’s CLIP!📎 CLIP models learn from language supervision, and are capable of strong zero-shot performance at various vision tasks (arxiv.org/abs/2103.00020) Our reproduction can be found at github.com/mlfoundations/…
    arXiv logo
    arxiv.org
    Learning Transferable Visual Models From Natural Language Supervision
    State-of-the-art computer vision systems are trained to predict a fixed set of predetermined object categories. This restricted form of supervision limits their generality and usability since...
  • user avatar
    Gabriel Ilharco
    @gabriel_ilharco
    Feb 23, 2021
    Instead of a single neural network, why not train lines, curves and simplexes in parameter space? Fantastic work by @Mitchnw et al. exploring how this idea can lead to more accurate and robust models: arxiv.org/abs/2102.10472
    Image
  • user avatar
    Gabriel Ilharco
    @gabriel_ilharco
    Nov 7, 2023
    Another breakthrough in CLIP models, powered by better datasets. Great job @Vaishaal, @AlexFang26 and team! Paper: arxiv.org/abs/2309.17425
    Image
    55K055K
  • user avatar
    Gabriel Ilharco
    @gabriel_ilharco
    Oct 12, 2020
    I've been seeing a lot of talk around the recent Vision Transformer (ViT) paper, so I thought I'd highlight some of my favorite previous work on self-attention and transformers in computer vision! Link to ViT: openreview.net/pdf?id=YicbFdN… (thread 👇)
  • user avatar
    Gabriel Ilharco
    @gabriel_ilharco
    Aug 12, 2022
    The year is 2032. A model was trained on all images, videos and text on the web, using over 100 yottaFLOPs. It still thinks this is an image of a dog. To fix models post-hoc, check out PAINT!🎨 📜 arxiv.org/abs/2208.05592 💻 github.com/mlfoundations/… 🌐 model-patching.github.io
    Image
  • user avatar
    Gabriel Ilharco
    @gabriel_ilharco
    May 1, 2020
    Vision plays a central role in shaping the meaning of concrete words like "apple" or "banana". Yet, most of today's NLP models learn representations of these concepts from text-only. Can such representations share similarities with the visual world? gabrielilharco.com/publications/p… 1/n
    Image
  • user avatar
    Gabriel Ilharco
    @gabriel_ilharco
    Feb 8, 2021
    Forget about messy vision backbones inside vision+language models? Check out ViLT, a cool work by Kim et al., extending Vision Transformers to multimodal domains. Link: arxiv.org/pdf/2102.03334…
    Image
  • user avatar
    Gabriel Ilharco
    @gabriel_ilharco
    Mar 5, 2021
    CLIP has spoken.
    Image
    Image
  • user avatar
    Gabriel Ilharco
    @gabriel_ilharco
    Sep 22, 2019
    Can machines learn language from grounded, untranscribed speech? I don't know, but we're making fast progress! Paper: arxiv.org/abs/1909.08782 Thread below (1/n)
    Image
    GIF