Technology You Will Work With:
- Large language models consumed through APIs (Alibaba Qwen / Model Studio, Anthropic Claude, OpenAI and comparable providers).
- Prompt engineering, system instructions, structured (JSON) outputs and tool / function calling.
- Retrieval-augmented generation: embeddings, chunking and vector search.
- Speech-to-text, text-to-speech and image and document recognition (OCR) services.
- Python or TypeScript for the scripts, integrations and evaluations you write.
- Evaluation, prompt versioning and observability tooling, with Git and GitHub for review and history.
Key Responsibilities:
- Design, write and iterate on the prompts and system instructions that drive the product's AI assistant, and version them the way code is versioned.
- Turn business rules and operating procedures into instructions, examples and constraints that a model follows reliably.
- Define and maintain the assistant's tools — what each one does, the arguments it takes and when the model should reach for it — working with engineers on the implementation.
- Prepare the knowledge the assistant reads: collect, clean, chunk and structure source material, and keep it current as the business changes.
- Build evaluation sets for every AI feature, run them on each change, and report accuracy, regressions and failure patterns with evidence rather than impressions.
- Integrate and tune speech recognition and image or document recognition so they hold up on real user input, including accented speech and poor-quality photographs.
- Investigate reported AI failures — a wrong answer, an invented fact, a missed or mis-called tool — and trace each one to its cause in the prompt, the retrieved data or the tool definition.
- Monitor token usage, latency and cost per request, and propose changes that reduce them without weakening quality.
- Test the assistant across the languages it serves and confirm it answers correctly in each of them.
- Apply responsible-use practice: keep sensitive data out of prompts and logs, and flag outputs that should never be produced.
- Write clear documentation of what the assistant can and cannot do, and update the in-app user manual whenever your change alters what a user sees or gets.
Requirements:
- Bachelor's Degree in Computer Science, Information Technology, Artificial Intelligence, Data Science, Engineering, or a related field.
- 1–3 years of experience building with large language models or other AI services. Internships and substantial personal or open-source AI projects will be considered.
- Demonstrable prompt engineering experience with a commercial LLM, and the judgement to explain why one prompt outperforms another.
- Working understanding of how these models behave: tokens and context windows, temperature, embeddings, structured output, and the common causes of hallucination.
- Comfortable reading and writing code to call APIs and process JSON — Python or JavaScript / TypeScript is sufficient; you do not need to be a full-stack engineer.
- Able to design test cases and evaluate model output objectively against them.
- Familiarity with structured data: JSON, CSV and basic SQL.
- Comfortable with Git and a pull-request based workflow.
- Strong analytical, problem-solving and communication skills, with attention to detail.
- Excellent command of written and spoken English.