|
Linguist II
Description: Onsite - 3 days in one of preferred offices: Sunnyvale CA, New York, Burlingame Washington, or Redmond WA.
Bachelor's degree required. We are looking for a linguist to help develop language components for AI-powered products, including large language models (LLMs) and voice-enabled technologies. We are seeking candidates with solid linguistic data analysis skills, programming familiarity, and language technology experience to contribute to data collection, synthetic data generation, and annotation tasks in support of LLM/AI training, evaluation, alignment, and AI agent development. Job Responsibilities Apply linguistic expertise in syntax, semantics, pragmatics, and sociolinguistics to support LLM and generative AI systems Collaborate with linguists, data operations teams, and ML engineers on data collection, curation, annotation, and localization efforts for model training and fine-tuning Contribute to the development and maintenance of annotation schemas and guidelines for LLM training data (e.g., instruction-tuning, preference labeling, RLHF) Evaluate and quality-check datasets used for pre-training, fine-tuning, and alignment of language models Support the development of programmatic methods for generating synthetic annotated data at scale Assist in model evaluation efforts including prompt-based testing, red-teaming, and linguistic error analysis Participate in experiments to assess data quality, annotation consistency, and downstream model performance Basic Qualifications Bachelor's degree in Linguistics, Computational Linguistics, Computer Science, Speech Science, or related field 1+ years of experience in Linguistics, Language Technologies, NLP, or AI/ML data operations (or equivalent) Knowledge of syntax, semantics, pragmatics, sociolinguistics, corpus linguistics, and other areas of linguistics Familiarity with Large Language Models (LLMs), their applications and data practices (training data, evaluation, prompting, fine-tuning) Exposure to LLM evaluation methodologies (human evaluation, automated metrics, adversarial testing) Experience working with semantic ontologies, taxonomies, or intent/slot frameworks Proficiency using AI Agents/Chatbots Experience with database queries and data analysis processes (SQL, spreadsheets, R, Unix, or others) Experience working with speech and text data in multiple languages Comfortable working in a fast-paced, highly collaborative environment with evolving priorities Preferred Qualifications Native or near-native fluency in English and at least one additional language Master's degree in Linguistics, Computational Linguistics, Language Technologies, or a related field Familiarity with machine learning frameworks, NLP libraries, and tools (e.g., Hugging Face, spaCy, NLTK, PyTorch) Exposure to statistical language modeling or training data pipelines Strong organizational skills and attention to detail Pursuant to the California Fair Chance Act, Los Angeles County Fair Chance Ordinance for Employers, Los Angeles Fair Chance Initiative for Hiring Ordinance, and San Francisco Fair Chance Ordinance, qualified applicants will be considered for assignment with arrest and conviction records. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness, meet client expectations, standards, and accompanying requirements, and safeguard business operations and company reputation. Custom Fields: Name: Story Behind the Need – Business Group & Key Projects Value: We are looking for a linguist to help develop language components for AI-powered products, including large language models (LLMs) and voice-enabled technologies. We are seeking candidates with solid linguistic data analysis skills, programming familiarity, and language technology experience to contribute to data collection, synthetic data generation, and annotation tasks in support of LLM/AI training, evaluation, alignment, and AI agent development. Name: If yes, approximately how often might CW’s engaged by your team encounter graphic or objectionable content (imagery, video, or written) in the course of their work with your team? Value: None Name: Compelling Story & Candidate Value Proposition Value: Job Responsibilities Apply linguistic expertise in syntax, semantics, pragmatics, and sociolinguistics to support LLM and generative AI systems Collaborate with linguists, data operations teams, and ML engineers on data collection, curation, annotation, and localization efforts for model training and fine-tuning Contribute to the development and maintenance of annotation schemas and guidelines for LLM training data (e.g., instruction-tuning, preference labeling, RLHF) Evaluate and quality-check datasets used for pre-training, fine-tuning, and alignment of language models Support the development of programmatic methods for generating synthetic annotated data at scale Assist in model evaluation efforts including prompt-based testing, red-teaming, and linguistic error analysis Participate in experiments to assess data quality, annotation consistency, and downstream model performance Name: Team Name Value: None Name: Top 3 must-have HARD skills Value: 1+ years of experience in Linguistics, Language Technologies, NLP, or AI/ML data operations (or equivalent) Knowledge of syntax, semantics, pragmatics, sociolinguistics, corpus linguistics, and other areas of linguistics Familiarity with Large Language Models (LLMs), their applications and data practices (training data, evaluation, prompting, fine-tuning) Exposure to LLM evaluation methodologies (human evaluation, automated metrics, adversarial testing) Experience working with semantic ontologies, taxonomies, or intent/slot frameworks Proficiency using AI Agents/Chatbots Experience with database queries and data analysis processes (SQL, spreadsheets, R, Unix, or others) Name: Typical Day in the Role Value: Job Responsibilities Apply linguistic expertise in syntax, semantics, pragmatics, and sociolinguistics to support LLM and generative AI systems Collaborate with linguists, data operations teams, and ML engineers on data collection, curation, annotation, and localization efforts for model training and fine-tuning Contribute to the development and maintenance of annotation schemas and guidelines for LLM training data (e.g., instruction-tuning, preference labeling, RLHF) Evaluate and quality-check datasets used for pre-training, fine-tuning, and alignment of language models Support the development of programmatic methods for generating synthetic annotated data at scale Assist in model evaluation efforts including prompt-based testing, red-teaming, and linguistic error analysis Participate in experiments to assess data quality, annotation consistency, and downstream model performance Name: Does this role deal with graphic and/or objectionable content to any extent? Value: No Name: Good to have skills Value: Preferred Qualifications Native or near-native fluency in English and at least one additional language Master's degree in Linguistics, Computational Linguistics, Language Technologies, or a related field Familiarity with machine learning frameworks, NLP libraries, and tools (e.g., Hugging Face, spaCy, NLTK, PyTorch) Exposure to statistical language modeling or training data pipelines Strong organizational skills and attention to detail Experience working with speech and text data in multiple languages Comfortable working in a fast-paced, highly collaborative environment with evolving priorities Name: How will performance be measured Value: Contribute to the development and maintenance of annotation schemas and guidelines for LLM training data (e.g., instruction-tuning, preference labeling, RLHF) | ||||