machinelearning
π Built a House Price Prediction Model using Python & Machine Learning
I built a House Price Prediction project as part of my Data Science Internship at Oasis Infobyte . π Project Overview The goal of this project is to predict house prices based on different property-related features using Machine Learning . π Project Workflow Dataset β Data Cleaning β EDA β Feature Selection β Model Training β Evaluation β Price Prediction π§Ή Data Preparation Loaded and explored the dataset using Pandas Handled missing and inconsistent data Prepared the dataset for machine learning π Exploratory Data Analysis Analyzed relationships between different features Created visualizations to identify patterns and trends Studied factors affecting house prices π€ Machine Learning Selected relevant features Split the dataset into training and testing sets Trained a regression model Evaluated the model using appropriate performance metrics π― Prediction The trained model takes property-related features as input and generates an estimated house price . π οΈ Tech Stack Python Pandas NumPy Matplotlib Scikit-learn Jupyter Notebook π‘ Key Learning This project gave me hands-on experience with the complete machine learning workflow β from data preprocessing and visualization to model training, evaluation, and prediction . π GitHub Repository πhttps://github.com/pallavisagar07/OIBSIP/tree/ec73e94f0a9d2ea3a3fe3e65c1676d8c0af17219/DataAnalytics-L2-House-Price-Prediction-Linear-Regression Python MachineLearning DataScience DataAnalytics ScikitLearn OasisInfobyte ProjectShowcase
π Built a Basic RAG and an Advanced RAG System from Scratch β Without LangChain
To understand how RAG works internally, I first built a Basic RAG Pipeline: Basic RAG PDF β Extraction β Chunking β Embeddings β FAISS Search β Top-K Chunks β Prompt β LLM β Answer It works well for simple questions, but struggles with exact keywords, IDs, and complex queries. So I built an Advanced RAG Pipeline: Advanced RAG PDF β Extraction β Chunking β Embeddings β Dense + Sparse Index β Query Understanding β Dense Retrieval + Sparse Retrieval β Reciprocal Rank Fusion (RRF) β Cross-Encoder Reranking β Noise Filtering β Context Fusion β Prompt β LLM β Answer Key Improvements β Semantic Search (FAISS) β Exact Keyword Search (BM25) β Hybrid Retrieval β Cross-Encoder Reranking β Duplicate Removal β Better Context Selection β Reduced Hallucinations One of the biggest lessons from this project: A better RAG system isn't always about using a better LLM. It's often about building a better retrieval pipeline. Tech Stack: Python β’ Streamlit β’ FAISS β’ Sentence Transformers β’ BM25 β’ Cross-Encoder Reranker β’ Groq RAG GenAI LLM AIEngineering MachineLearning Python FAISS BuildInPublic
HeritageAI
Building HeritageAI ποΈπ€ β a voice-controlled PowerPoint assistant that lets presenters navigate slides hands-free using natural voice commands. Currently improving it with smarter speech interpretation and AI-powered presentation features. π π https://bit.ly/4Aj9LP0 AI MachineLearning VoiceAI Python HeritageAI