EchoTwin AI

EchoTwin AI

Vision Language Model Engineer

San Francisco Office

Sponsorship not specifiedDetected 319 days ago
PythonAWSGCPAzureCloud PlatformsMachine LearningDeep LearningPyTorchNLPComputer VisionResearchCommunicationProblem Solving

About the role

  • Company Overview EchoTwin AI is pioneering AI-driven infrastructure intelligence, redefining how cities are managed.
  • More than "smart cities," EchoTwin is advancing the era of cognizant cities-urban environments with the awareness to see, think, and act on challenges in real time.
  • EchoTwin AI is not responsible for any fees related to unsolicited resumes.

Responsibilities

  • You will work closely with cross-functional teams to build models that power applications such as image captioning, visual question answering, and multimodal AI at the edge.
  • Collaborate with data scientists and software engineers to integrate models into production systems.
  • Optimize model performance for accuracy, latency, and scalability in real-world applications.
  • Document model development processes and present findings to technical and non-technical stakeholders.
  • By enabling municipalities to proactively monitor, predict, and resolve issues, EchoTwin helps build resilient, self-healing, and sustainable urban ecosystems.
  • At EchoTwin AI, we value excellence regardless of background and are committed to building a team that reflects the communities we serve.

Requirements

  • Proficiency in Python and relevant ML libraries (e.g., Hugging Face, OpenCV, Transformers).
  • Experience with large-scale model training and optimization (e.g., distributed training, quantization).
  • Strong understanding of neural network architectures (e.g., CNNs, Transformers, CLIP, or similar).
  • Experience with multimodal datasets and preprocessing techniques for images and text.
  • Familiarity with cloud platforms (e.g., AWS, GCP, Azure) and model deployment workflows.
  • Strong problem-solving skills and ability to work in a fast-paced, collaborative environment.
  • Bachelor's, Master's or Ph.D. in Computer Science, Machine Learning, Artificial Intelligence, or a related field (or equivalent experience).
  • 3+ years of experience in machine learning, with a focus on vision-language models or multimodal AI.
  • Hands-on experience with deep learning frameworks such as PyTorch or TensorFlow.
  • Proven track record of building and deploying computer vision and/or NLP models.
  • Excellent communication skills to explain complex technical concepts to diverse audiences.
  • Benefits and Perks
  • There are endless learning and development opportunities from a highly diverse and talented peer group, including experts in various fields, including Computer Vision, GenAI, Digital Twin, Government Contracting, Systems and Device Engineering, Operations, Communications, and more!
  • Options for medical, dental, and vision coverage for employees and dependents (for US employees)

Benefits

  • Flexible Spending Account (FSA) and Dependent Care Flexible Spending Account (DCFSA)
  • Profit sharing
  • As a Vision Language Model Engineer, you will design, develop, and optimize advanced vision-language models that integrate visual and textual data to enable intelligent systems.
  • Design and implement state-of-the-art vision-language models using deep learning frameworks.
  • Develop and fine-tune models that combine computer vision and natural language processing for tasks like image captioning, visual question answering, and text-to-image generation.
  • Stay up-to-date with the latest research in vision-language models and incorporate advancements into projects.

Equal opportunity

  • Equal Opportunity Employer.

This listing is sourced directly from EchoTwin AI's careers page and normalized into a canonical job model.