

Search by job, company or skills

Role Description▍ About the Role
We're hiring a QA Engineer to test our Agent-as-a-Service conversational AI platform for Arabic-speaking customers, focused on Kuwaiti and Najdi dialects. This isn't traditional QA — expected output isn't an exact string match. You'll judge whether the agent understood intent, replied in natural correct dialect, followed business rules, held context, retrieved accurate info, and stayed safe, across real conversations where correctness is linguistic and contextual, not binary.
Arabic LLM research consistently shows MSA fluency does not predict dialect fluency — dialect, culture, code-switching, reasoning, and safety fail independently and must be tested as separate dimensions, not assumed to travel together.
▍ Responsibilities
1. Dialect & Linguistic QA
Test Kuwaiti/Najdi responses for: dialect authenticity (no unprompted MSA drift), grammar, vocabulary, tone/register, naturalness, spelling/punctuation, Arabic-English code-switching, numeral style, Arabizi. Flag answers that are technically correct but don't read as native.
2. Conversational QA
Test full multi-turn conversations, not isolated prompts: context retention, pronoun resolution, corrections, topic switching, long threads, angry/emotional customers, typos, voice-transcribed Arabic, mixed Arabic/English. Confirm the agent never loses context or contradicts itself.
3. Intent & Instruction Following
Verify correct handling of booking/reschedule/cancel, pricing/service questions, complaints, escalation triggers, multi-intent and implicit/ambiguous messages. Verify business rules hold even under user attempts to override them.
4. RAG & Knowledge QA
Check answers against source knowledge; catch hallucinations, missing/conflicting/outdated info, and ambiguous-document confusion. Distinguish a correct+grounded answer from a correct-looking but unsupported one. Include dialectal queries against MSA knowledge sources.
5. Cultural & Regional QA
Check local terminology, Kuwaiti/Saudi conventions, names/places, dates, currency, formality register, local expressions/idioms — as its own pass, not a grammar check.
6. Safety & Adversarial Testing
Run prompt injection, jailbreaks, instruction override, system-prompt extraction, abusive/harmful input, social engineering, malicious links — in Kuwaiti/Najdi Arabic, not just English. Arabic-specific safety blind spots don't show up in English-only testing.
7. Regression Suite
Every confirmed bug (dialect, hallucination, context, intent, RAG, safety, prompt-injection) becomes a permanent test case. Run full suite after every prompt/model/pipeline change, diff against prior version.
▍ Example of the Bar
User: ممكن تحجزلي موعد باجر العصر؟
Not is the answer correct but: did it parse باجر / العصر correctly, keep dialect, sound native, ask for missing info instead of inventing availability, follow the booking workflow, and hold that context next turn
▍ Evaluation Dimensions
Each conversation is scored across:
Dimension
Question
Task Success
Did it accomplish the user's goal
Intent Accuracy
Did it understand the request
Dialect Fidelity
Consistent Kuwaiti/Najdi, no drift
Naturalness
Native-sounding
Context Retention
Remembers prior turns
Factuality
Information correct
Groundedness
Supported by knowledge base
Instruction Following
Obeys system/business rules
Cultural Appropriateness
Regionally correct
Safety
Handles adversarial input
Robustness
Survives typos/slang/ambiguity
Escalation
Hands off correctly when required
▍ Required
• Native/near-native Arabic, strong Kuwaiti and/or Najdi dialect fluency
• Experience testing chatbots, conversational systems, or LLM applications
• Strong exploratory + adversarial test design
• Understands LLM failure modes; judges by semantic correctness, not string match
• Good written English for bug reports; Jira/Linear or equivalent
▍ Strongly Preferred
• LLM evaluation / LLM-as-a-Judge experience
• RAG system testing
• Prompt engineering understanding
• Hands-on with OpenAI/Anthropic/Gemini or similar
• API testing (Postman/cURL); basic Python for test automation
• Experience building regression/eval datasets; Arabic NLP background
Qualifications
Job ID: 152477333