Job Details
Research Engineer, Pretraining
About KiteFishAI
KiteFishAI is an AI research lab focused on building efficient, domain-specialized and sovereign Large Language Models (LLMs) and Small Language Models (SLMs).
We work across the full model development stack — from data and model architecture to pretraining, post-training, evaluation, inference and deployment. Our goal is to build models that are not only capable, but efficient, reliable and useful for real-world applications.
We are a research-first team. We believe that building strong models requires understanding what happens underneath the APIs and frameworks — from data quality and tokenization to optimization, distributed training and model behavior.
We are looking for a Research Engineer — LLM Pretraining & Systems to work closely with our research team on building and improving the next generation of KiteFishAI models.
What You’ll Work On
- Research and implement improvements across LLM architecture, pretraining, optimization, data and evaluation
- Build and experiment with Transformer-based language models from the ground up
- Design and execute controlled experiments to understand model behavior and training dynamics
- Develop data pipelines for collecting, cleaning, filtering, deduplicating and preparing large-scale training datasets
- Experiment with tokenization, data mixtures, sequence lengths and sampling strategies
- Work on training efficiency, memory optimization and GPU utilization
- Implement and optimize distributed training across multiple GPUs and machines
- Investigate training instability, convergence, scaling and model quality
- Evaluate models using both standard benchmarks and internally designed evaluations
- Read and reproduce ideas from recent research papers and determine whether they improve our models
- Build internal tools for training, evaluation, experimentation and model analysis
- Contribute across the stack — from low-level GPU and PyTorch optimizations to high-level model design
- Work closely with the team to take research ideas from paper → experiment → implementation → production model
What We’re Looking For
- Strong programming and software engineering fundamentals
- Excellent Python skills and hands-on experience with PyTorch or another deep learning framework
- Strong understanding of machine learning and deep learning fundamentals
- Good understanding of Transformer architectures and how modern language models work
- Ability to understand research papers and translate them into working implementations
- Strong problem-solving and debugging skills
- Curiosity to understand systems at a deeper level rather than treating ML frameworks as black boxes
- Ability to design experiments, analyze results and draw conclusions from data
- Comfortable working in an environment where research directions can change quickly
- Strong ownership and ability to independently take a problem from idea to implementation
Preferred Experience
You do not need prior experience training a large language model from scratch.
Experience in any of the following areas is valuable:
- Training or fine-tuning Transformer-based models
- Building ML or deep learning systems using PyTorch
- Distributed computing or distributed machine learning
- GPU programming, CUDA or performance optimization
- Large-scale data processing or ETL pipelines
- NLP or language modeling
- Experience with GPUs, Kubernetes or cloud-based ML infrastructure
- Understanding of optimizers such as Adam/AdamW, learning-rate schedules and training dynamics
- Experience with mixed-precision training, gradient accumulation, checkpointing or memory optimization
- Familiarity with technologies such as FSDP, DeepSpeed, Megatron-LM or similar distributed training systems
- Experience reproducing results from research papers
- Contributions to open-source ML/AI projects
- Research experience in machine learning, NLP or related areas
You’ll Thrive at KiteFishAI If You
- Want to understand how LLMs actually work, not just how to use them
- Enjoy reading papers and implementing ideas from first principles
- Are comfortable experimenting, failing and trying again
- Care about measurement and scientific rigor
- Can move between research and engineering without treating them as separate disciplines
- Are willing to work on problems outside your immediate area of expertise
- Prefer building systems yourself over relying entirely on existing abstractions
- Have a strong bias toward learning and execution
- Want to work on models where you can see the connection between an idea, an experiment and the final model
Sample Projects
Depending on your background, you may work on projects such as:
- Implementing and benchmarking a new Transformer architecture
- Improving training throughput and GPU utilization
- Experimenting with different optimizers and learning-rate schedules
- Studying the effect of different data mixtures on model capabilities
- Building large-scale data filtering and deduplication pipelines
- Investigating why a model's training loss stops improving
- Reproducing a recent LLM research paper
- Scaling model training from a single GPU to multi-node clusters
- Implementing distributed training and fault-tolerant checkpointing
- Improving inference efficiency and memory usage
- Designing evaluation suites for domain-specific language models
- Analyzing model failures and developing experiments to understand their causes
The Ideal Candidate
We care more about fundamentals, curiosity and ability to learn than whether you have previously trained a 1B or 7B model or a simple ML model.
If you have built a Transformer yourself, implemented a paper in PyTorch, trained models from scratch, understand why AdamW works, can debug a GPU memory issue, or can look at a training curve and reason about what is happening — we want to hear from you.
You don't need to have trained an LLM at scale before. You need to be capable of learning how to do it.