Named Entity Recognition (NER) is used to identify different types of entities in text, such as people, places, and organizations. One of the most difficult parts of this project was the limited availability of rich datasets for Roman Urdu–English code-mixed text. Roman Urdu–English code-mixed text means that different languages, such as English and Roman Urdu, can be used within the same sentence. This makes NER challenging because Roman Urdu can be written in many different ways. For example, one person may write a name as “Mehwish,” while another person may write it as “Mevish.” People can also spell Roman Urdu words differently according to their own writing style. Because of these spelling variations and the limited amount of available data, it is difficult for a model to learn all possible variations and make reliable predictions on different types of code-mixed text. The Problem with Code-Mixed NER The main challenge in this project was the many variations in Roman Urdu spelling. Different people can use their own way of writing the same Roman Urdu word, which makes it difficult for a model to learn consistent patterns. Another challenge is that Roman Urdu and English can appear in the same sentence. The language can switch within a single sentence, so the model needs to learn patterns from the context rather than simply memorizing names. For example, it needs to learn whether a word or phrase represents a person, location, or organization based on how it is used in the sentence. The hardest part was the limited availability of Roman Urdu–English code-mixed datasets. With a relatively limited source of training data, it becomes more difficult for the model to learn the different spelling and language patterns that occur in real-world text. My project therefore focuses on recognizing three entity types — Person, Location,