[{"content":"TL;DR: Digital Dhaba is a one-page weekly newsletter: code gathers and ranks the week\u0026rsquo;s stories, Claude writes the digest only from those stories, and every item links to the original source, so you can check it in one click.\nThe problem Tech and AI news never stops, and it is spread across dozens of sites, newsletters and forums. Most people can\u0026rsquo;t follow it all, and much of it is written for specialists.\nWhat I built Tech Digest is a one-page weekly newsletter, published under the name Digital Dhaba, that tells a non-specialist what mattered this week in tech, AI and LLMs. Every item links to the original source. Each issue has:\nTop Stories: the 5 most important stories of the week. Topic sections: the rest of the news, grouped by subject. Quick Hits: 5–8 one-line stories worth knowing. Try This: one practical AI tip. Worth Your Weekend: 2–3 essays, talks, papers or books worth slowing down for. How one issue gets built One script, run_digest.sh, runs four steps in order:\nGather: fetch.py pulls the week\u0026rsquo;s stories from Hacker News, research paper sites, tech news sites and other newsletters. It removes duplicates, sorts stories into sections and ranks them by how much attention they got. Write: Claude gets those stories plus the instructions and writes the digest as Markdown. Convert: a Node script turns the Markdown into structured data (digest.json). Design: another Node script renders that into an email-safe HTML page (newsletter.html). All three files land in issues/YYYY-MM-DD/.\nDecisions I took Claude writes with no tools and no web access. It can only use the stories it was handed, so it can\u0026rsquo;t invent news. Ranking is done by code, not by AI. The same input always gives the same ranking, and only the writing step costs tokens. Weekend picks come from a hand-checked shelf (weekend_picks.yaml), with summaries verified in advance. Tips come fresh from The Neuron\u0026rsquo;s last 7 days, with tips.yaml (now 9 tips) as the fallback when there is no good fresh one. Nothing repeats. The script reads past issues to see what has already run, so there is no history file to maintain. Test issues don\u0026rsquo;t count. no_repeat_since: 2026-10-08 means picks used before launch can run again for real readers. A GitHub Action runs it, not a laptop or a cloud machine I have to keep switched on, and it runs the exact same script every week. Cost is about $0.93 per issue at API prices; on my Claude subscription it counts against plan usage instead. How it gets published Every Thursday, GitHub starts a fresh temporary machine and runs run_digest.sh (weekly-digest.yml). If the digest comes out empty or broken, the run fails and nothing is published. Otherwise it commits the new issue folder to the Digital Dhaba repo. It then starts the second workflow, publish-to-blog.yml, which copies the issue into my blog repo so it appears on blogsbykush.com. The one thing The rule I care most about: it only writes from what it collected. Nothing from the model\u0026rsquo;s memory, nothing made up to fill space. Every item links back, so you can check it yourself in one click. An AI-written newsletter is only useful if you can verify it. That link is the whole trust model.\nRead the latest issue, or get it every Thursday: subscribe.\nPart of my Build Log.\n","permalink":"https://blogsbykush.com/build-log/digital-dhaba-ai-newsletter-that-cannot-make-up-news/","summary":"A weekly tech digest where code picks the stories, Claude only writes from them, and every item links to its source.","title":"Digital Dhaba: how I built an AI-written tech newsletter that can't make up news"},{"content":" ‹ 1 / 4 › Download as PDF Read this comic as text Slide 1: Remember Antakshari?\nTeacher: Remember playing Antakshari at family functions? The next song starts with the last letter of the previous one.\nStudent: Ha, yes! Someone sings \u0026ldquo;…gaata hoon,\u0026rdquo; and the next person scrambles for a song starting with \u0026ldquo;N.\u0026rdquo;\nTeacher: Notice, nobody plans the whole session in advance. Each person looks at the last sound and picks the best next song they know.\nStudent 2: True. Nobody says \u0026ldquo;I have a 10-song strategy.\u0026rdquo; It\u0026rsquo;s one move at a time.\nSlide 2: One move at a time\nTeacher: That one move, picking the next best song based on what came before, is exactly what an LLM does. Only its unit isn\u0026rsquo;t a song, it\u0026rsquo;s a word. Actually, even smaller: a piece of a word, called a token.\nStudent: Wait, so it\u0026rsquo;s not \u0026ldquo;thinking\u0026rdquo; of the full answer first and then writing it out?\nTeacher: No plan, no destination. It looks at everything said so far, asks \u0026ldquo;what usually comes next?\u0026rdquo;, and picks that.\nStudent 2: So if I ask it to write an email, it\u0026rsquo;s not thinking \u0026ldquo;here\u0026rsquo;s my conclusion, let me build up to it\u0026rdquo;?\nSlide 3: A superhuman Antakshari player\nTeacher: Exactly. It\u0026rsquo;s a very, very good Antakshari player who has heard almost every song ever sung, and instinctively knows what fits next. One word at a time, until the email is done.\nStudent: That\u0026rsquo;s humbling. It feels intelligent because each move is impressively good, not because there\u0026rsquo;s a grand plan.\nTeacher: So in your own words, what\u0026rsquo;s an LLM, really?\nStudent 2: It\u0026rsquo;s a next-word prediction engine. Not a chess player thinking ahead, just Antakshari at superhuman level, one word after another.\nSlide 4: The whole trick\nTeacher: That\u0026rsquo;s the whole trick. No goal, no plan, just really, really good next-move prediction, repeated until it looks like thought.\nTakeaway: An LLM doesn\u0026rsquo;t plan answers. It predicts the next token, again and again.\nThe concept in plain words A large language model (LLM), the technology behind ChatGPT, Claude and Gemini, does one thing: given some text, it predicts the piece of text most likely to come next. That piece is a token, usually a word or part of a word.\nWhen you ask a question, it predicts one token, adds it to the text, and predicts the next one using everything written so far. It repeats this hundreds of times until the answer is done. That\u0026rsquo;s the Antakshari loop: look at what came before, pick the best next move, repeat.\nIts moves are good because it was trained on an enormous amount of text. Like the player who has \u0026ldquo;heard almost every song\u0026rdquo;, it has seen so many patterns that its next guess is usually right.\nOne difference from Antakshari: an LLM doesn\u0026rsquo;t just look at the last word, it rereads the whole conversation every time. That\u0026rsquo;s why the context you give it matters so much.\n","permalink":"https://blogsbykush.com/concept-breakdown/how-llms-think/","summary":"Hint: it\u0026rsquo;s just Antakshari. How a large language model writes an answer, one token at a time.","title":"How LLMs Think"},{"content":"TL;DR: Training an AI model for one skill also shifts its choices slightly on unrelated questions, which makes me think hiding a model\u0026rsquo;s reasoning is one layer of protection against copying, not the whole wall.\nThe source A 2026 research paper (a preprint, not yet peer reviewed) by Ziyang Zhang and colleagues, listed on Hugging Face under Peking University: Post-Training Leaves Behavioral Shadows on Unrelated Decisions.\nThe idea, in my words When you take a public base model and post-train it for one skill, such as coding, the change doesn\u0026rsquo;t stay inside that skill. The model\u0026rsquo;s choices also shift slightly on questions that have nothing to do with it. The paper suggests those small shifts can be picked up by another model that starts from the same base, and that they carry part of the skill with them. The authors call it a low-bandwidth channel: it works between models that share a base, and the gains they report are modest.\nWhat I took from it Hiding a model\u0026rsquo;s reasoning may be one layer of protection against copying, not the full boundary. Small signals may leak through off-topic answers, at least between models that share a base. For anyone building on the same few open models, that\u0026rsquo;s a reminder that \u0026ldquo;what a model learned, and from whom\u0026rdquo; is harder to pin down than it looks.\nWhat I\u0026rsquo;m still unsure about Does this matter between models that don\u0026rsquo;t share a base? The paper\u0026rsquo;s results are on small models, so I don\u0026rsquo;t know yet how much of this carries over to the frontier models people actually use.\nFurther reading The paper on arXiv: the method and the results. The paper\u0026rsquo;s page on Hugging Face: community discussion. Subliminal Learning (Anthropic): the 2025 work on traits moving between models through ordinary-looking data, which this paper builds on. Part of my Learning Notes. Found via Digital Dhaba.\n","permalink":"https://blogsbykush.com/learning-notes/training-an-ai-for-one-skill-changes-its-other-answers/","summary":"Post-training a model for one skill nudges its answers on unrelated questions, and that changes how I think about protecting models from copying.","title":"Training an AI for one skill quietly changes its other answers"},{"content":"Top stories Gemini 4 Argon: Google\u0026rsquo;s new frontier model drew the week\u0026rsquo;s biggest AI thread (1,688 HN points), though access is limited to government users and trusted cyber defenders for now. GPT 6.1 Sol: OpenAI says it delivers near-Astra intelligence at a fifth of the price, extending the model price war. Claude Sonnet 5.5: Anthropic says it runs 30%+ faster and costs up to 30% less for most work, at the same price as Sonnet 5. When did Google get so weird?: This was the most-discussed non-AI post of the week (2,009 HN points, 1,116 comments). Hacks of 2 federal agencies in a month have spilled a bonanza of sensitive data: The Pentagon is notifying more than 2 million current and former service members that their personnel records were stolen. ","permalink":"https://blogsbykush.com/tech-digest/2026-10-03/","summary":"The last 7 days in tech, AI \u0026amp; LLMs","title":"Digital Dhaba — October 3, 2026"},{"content":" Think of Statistics as the \u0026ldquo;mathematics of chai-making\u0026rdquo; - you need to know the right proportions of milk, water, sugar, and tea leaves to make the perfect cup, right? Similarly, Machine Learning without Statistics is like trying to make chai blindfolded - you might get lucky once, but you won\u0026rsquo;t know why it worked or how to repeat it!\nStatistics is literally the backbone of ML. When Netflix recommends \u0026ldquo;Sacred Games\u0026rdquo; after you watched \u0026ldquo;Mirzapur,\u0026rdquo; it\u0026rsquo;s using statistical patterns from millions of users. When your phone\u0026rsquo;s autocorrect knows you meant \u0026ldquo;biryani\u0026rdquo; not \u0026ldquo;biriyani,\u0026rdquo; that\u0026rsquo;s statistics at work! It helps ML understand patterns, make predictions, and most importantly, tell us how confident we are about those predictions.\nRemember: Machine Learning is just Statistics on steroids with a fancy computer doing the heavy lifting!\nMachine Learning is just Statistics on steroids with a fancy computer doing the heavy lifting\n","permalink":"https://blogsbykush.com/concept-breakdown/statistics-in-ml/","summary":"Machine Learning is just Statistics on steroids with a fancy computer doing the heavy lifting","title":"Statistics in ML"},{"content":" Just wrapped up my latest ML comic on Clustering - where data points find their natural friend groups! After exploring how models balance complexity (Regularization) and navigate the learning sweet spot (Bias-Variance), this one shows how algorithms can discover hidden patterns without being told what to look for.\nSwipe through the comic below to see how Teacher breaks down this unsupervised learning magic to Student - spoiler alert: it involves a fun analogy that\u0026rsquo;ll make clustering click instantly!\nDiscover the essentials of Clustering in our latest classroom conversation.Clustering - where data points find their natural friend groups!\n","permalink":"https://blogsbykush.com/concept-breakdown/what-is-clustering/","summary":"Discover the essentials of Clustering in a classroom comic. Clustering - where data points find their natural friend groups!","title":"What is Clustering"},{"content":" Dive into our latest classroom conversation where we demystify Bias and Variance! Good predictions need to be both accurate (hit the target) and consistent (hit the same spot). Check out our engaging comic strip to understand this crucial aspect step-by-step\nDiscover the essentials of Bias \u0026amp; Variance in our latest classroom conversation.Good predictions need to be both accurate (hit the target) and consistent (hit the same spot)!\n","permalink":"https://blogsbykush.com/concept-breakdown/what-is-bias-and-variance/","summary":"Discover the essentials of Bias \u0026amp; Variance in a classroom comic. Good predictions need to be both accurate (hit the target) and consistent (hit the same spot)!","title":"What is Bias and Variance"},{"content":" Dive into our latest classroom conversation where we demystify Regularization! Learn how transforming raw data into meaningful features can enhance your machine learning models. Check out our engaging comic strip to understand this crucial aspect step-by-step\nDiscover the essentials of Regularization in our latest classroom conversation.\n","permalink":"https://blogsbykush.com/concept-breakdown/what-is-regularization/","summary":"Discover the essentials of Regularization in a classroom comic.","title":"What is Regularization"},{"content":"Introduction Data versioning is a crucial aspect of machine learning workflows, ensuring that data changes are tracked and reproducible. Data Version Control (DVC) is a popular tool for this purpose, providing a way to version data similarly to how Git versions code.\nIn this blog, I want to share my experience of attempting to implement a data versioning pipeline using AWS Lambda and DVC. While my goal was to create a fully cloud-based solution, I encountered several challenges that made this approach less ideal.\nThis post is not about a successful implementation but rather a cautionary tale about the potential pitfalls of this combination.\nUse Case Data Version Control (DVC) lets you capture the versions of your data and models in Git commits while storing them on-premises or in cloud storage. It also provides a mechanism to switch between these different data contents, you can refer to the official DVC documentation here for more details on DVC.\nMy existing data versioning process involved uploading data to an AWS EC2 instance, where a scheduled cron job would execute a Python script containing DVC and Git commands to version the data and store hashed data files in an AWS S3 bucket.\nThe problem with this approach is the dependency on the AWS EC2 machine. I am using the machine only to upload the data and run a Python script. Maintaining EC2 instances can be costly, so I thought these basic functions could be done by AWS Lambda. AWS Lambda allows you to run code without provisioning or managing servers and is known for its serverless capabilities and ease of integration with other AWS services.\nThe new approach I envisioned was:\nUpload data to an S3 bucket. Trigger a Lambda function upon upload. The Lambda function would handle data versioning with DVC. Versioned files would be stored in Git, and hashed files would be stored back in the S3 bucket. Data Versioning Pipeline\nImplementation There are various ways to create AWS Lambda Functions: you can write them from scratch, upload them as a zip file, or create them from a Docker image. I tried all these methods and encountered several issues with each. For instance, when I tried to create a Lambda function from scratch, I couldn\u0026rsquo;t run DVC commands because they need to be pre-installed in AWS Lambda. I faced many such issues, but to keep this post short, I won\u0026rsquo;t list them all.\nTherefore, I decided to create the Lambda function using a Docker image. The image below explains the Lambda creation process:\nIn this approach, I created a Docker image using the following Dockerfile:\nFROM public.ecr.aws/lambda/python:3.11 # Installing GIT depedencies RUN yum update -y RUN yum install git -y # Copy requirements.txt COPY requirements.txt ${LAMBDA_TASK_ROOT} # Copy function code COPY lambda_function.py ${LAMBDA_TASK_ROOT} # Install the specified packages RUN pip3 install -r requirements.txt --target \u0026#34;${LAMBDA_TASK_ROOT}\u0026#34; # Set the CMD to your handler (could also be done as a parameter override outside of the Dockerfile) CMD [ \u0026#34;lambda_function.handler\u0026#34; ] Dockerfile\nIn this Dockerfile, I copied the requirements.txt file to install necessary dependencies like DVC and boto3 for handling S3 operations programmatically:\ndvc dvc[s3] s3fs boto3 Requirement File\nNext, I copied the following Python code, which serves as the main Lambda function to execute DVC commands and perform all operations for the data versioning process:\nimport json import boto3 import os import logging logger = logging.getLogger() logger.setLevel(logging.INFO) def lambda_handler(event, context): source_bucket =\u0026#39;mytestdata-upload\u0026#39; logger.info(\u0026#34;New files uploaded to the source bucket.\u0026#34;) key = event[\u0026#39;Records\u0026#39;][0][\u0026#39;s3\u0026#39;][\u0026#39;object\u0026#39;][\u0026#39;key\u0026#39;] logger.info(\u0026#34;Copying cloned GIT repo under the temp directory\u0026#34;) os.system(\u0026#34;cp -r /var/task/dvc-aws-lambda-s3-dataversioning-repo/ /tmp/\u0026#34;) try: logger.info(\u0026#34;Changing directory to download file from S3 bucket\u0026#34;) os.chdir(\u0026#34;/tmp/dvc-aws-lambda-s3-dataversioning-repo\u0026#34;) s3.Bucket(source_bucket).download_file(key, key) logger.info(\u0026#34;Running dvc add command to create hashed file\u0026#34;) os.system(\u0026#34;dvc add \u0026#34;+ key) logger.info(\u0026#34;Running git commands to add .dvc file to the GIT repo\u0026#34;) os.system(\u0026#34;git add .gitignore\u0026#34;) os.system(\u0026#34;git add *.dvc\u0026#34;) os.system(\u0026#39;git commit -m \u0026#34;adding data from AWS Lambda\u0026#34;\u0026#39;) os.system(\u0026#34;git push\u0026#34;) logger.info(\u0026#34;Running dvc push commands to pushed hashed file to the remote storage\u0026#34;) os.system(\u0026#34;dvc push\u0026#34;) except botocore.exceptions.ClientError as error: logger.error(\u0026#34;There was an error in versioning process\u0026#34;) print(\u0026#39;Error Message: {}\u0026#39;.format(error)) return { \u0026#39;statusCode\u0026#39;: 200, \u0026#39;body\u0026#39;: json.dumps(\u0026#39;Hello from Lambda!\u0026#39;) } Lambda Function Handler\nThe Lambda function handler is the method in your function code that processes events. When your function is invoked, Lambda runs the handler method. Your function runs until the handler returns a response, exits, or times out. For more details about this function, refer to this official document.\nTo download data from S3, you need to create roles and policies that allow AWS Lambda to access the S3 bucket. This will enable the Lambda function to download the data and perform the required actions.\nAfter creating the Lambda function, the following commands were executed to create and push Docker images and to create the Lambda function:\n# 1. Build Docker Image $ docker build --platform linux/amd64 -t docker-image:dvc-lambda . # 2. Login to ECR $ aws --region us-east-1 ecr get-login-password | docker login --username AWS --password-stdin 095******563.dkr.ecr.us-east-1.amazonaws.com # 3. Create the ECR repository $ aws --region us-east-1 ecr create-repository --repository-name dvc-lambda --image-scanning-configuration scanOnPush=true --image-tag-mutability MUTABLE # 4. Build the image - it might take a few minutes to complete this step $ docker tag docker-image:dvc-lambda 095******563.dkr.ecr.us-east-1.amazonaws.com/dvc-lambda:latest # 5. Push the image to ECR $ docker push 095******563.dkr.ecr.us-east-1.amazonaws.com/dvc-lambda:latest # 6. Create a Lambda function by following command or in console $ aws lambda create-function \\ --function-name dvc-lambda \\ --package-type Image \\ --code ImageUri=095******563.dkr.ecr.us-east-1.amazonaws.com/dvc-lambda-function:latest \\ --role arn:aws:iam::095******563:role/lambda-write-s3bucket # 7. Test the lambda function $ aws lambda invoke --function-name dvc-lambda response.json In the above mentioned steps, the Docker image creation and Lambda function are created successfully, but on invoking the Lambda function, the step that is executing the dvc add command is giving an error. I am covering the reason in the next section.\nLearnings During the implementation of this use case, I encountered several interesting challenges that can serve as valuable lessons for other use case implementations:\nLambda Timeout Issue: I was unable to execute basic DVC commands, like dvc --version, in AWS Lambda. I attempted to run the command via code in the lambda_handler function, which was packaged in a Docker image and deployed as an AWS Lambda function. After numerous attempts, I realized the issue was due to the default timeout value of Lambda. DVC commands in Lambda took more than 5 seconds to respond, while the default timeout configuration was 3 seconds. After increasing the timeout value, I was able to resolve this issue.\nLambda File System Access Issue:\nWrite to the /tmp Path: AWS Lambda\u0026rsquo;s file system is read-only except for the /tmp path. If you need to write to the file system in a Lambda function, ensure your code writes to a path inside the /tmp directory. Read-only Error with DVC Commands: Another issue I encountered was a read-only error when executing DVC commands in Lambda. The container image I used was public.ecr.aws/lambda/python:3.11, where the default directory is /var/task/. The lambda_handler.py function was copied to this directory when building the Docker image. After cloning the Git repo, I tried running the dvc add command, which created a corresponding DVC folder under /var/tmp/, containing some cache and files. This directory is crucial for running DVC operations. However, since all directories except /tmp are read-only, I encountered the following error during dvc add: ERROR: unexpected error - [Errno 30] Read-only file system: /var/tmp/dvc. To resolve this, I overrode the default path with a path inside the /tmp directory by running the command $ dvc config core.site_cache_dir .dvc/site_cache_dir. User Permissions and Directory Access: While debugging the Lambda function error, I discovered that the default directory /var/tasks/ is associated with the root user. However, directories created under /tmp are associated with the user sbx_user1051, belonging to group 990 (Docker usage group). Further investigation revealed that AWS Lambda functions have multiple sbx_users ranging from sbx_user1051 to sbx_user1176. DVC Add Command Timeout: The dvc add command was timing out. Despite multiple attempts, I could not identify the exact reason. Based on my inspection of folder permissions and access in AWS Lambda, I suspect this issue is related to user permissions. The site_cache_dir, which should be created during the dvc add execution, was not being created, likely due to permission issues. The /tmp directory in AWS Lambda is created with the user sbx_user1051, while in Docker, it is usually the root user. I could run DVC commands as the root user in Docker, but in Lambda, where I executed all commands, everything was under the user sbx_user1051. This discrepancy likely caused the dvc add command to fail, and the site_cache_dir was not created. To investigate, I created a non-root user with the following commands in the Dockerfile. However, after creating an image from this Dockerfile and logging into the container, I was unable to run DVC commands from the created sbx_user. DVC needs to be executed as the root user within the container. I\u0026rsquo;m not sure how to resolve this step yet. Last Word Writing this post was tricky for me. First, I have not yet successfully implemented the use case, and I usually write about things that I have fully completed. Second, I faced numerous issues, so the challenge was to present them in a clear and helpful manner without causing confusion. This is one reason this post has been pending for a long time. When I started implementing this use case, it seemed relatively straightforward. However, I quickly realized that creating a fully automated, cloud-based data versioning pipeline using AWS Lambda and DVC was not as simple as it appeared. I hope my learnings can be helpful to others.\nPlease let me know if anyone else has faced similar issues or has found a way to resolve them.\n","permalink":"https://blogsbykush.com/mlops/why-dvc-not-best-choice-with-awslambda-for-dataversioning/","summary":"Discover the challenges and learnings from attempting to implement a data versioning pipeline using AWS Lambda and DVC. This blog shares a real-world experience of navigating the pitfalls of combining these technologies for a cloud-based solution.","title":"Why AWS Lambda with DVC is not a best choice for Data Versioning Pipeline"},{"content":"\nDive into our latest classroom conversation where we demystify Feature Engineering! Learn how transforming raw data into meaningful features can enhance your machine learning models. Check out our engaging comic strip to understand this crucial aspect step-by-step\nDiscover the essentials of Feature Engineering in our latest classroom conversation. Learn how to convert raw data into meaningful features to boost your machine learning models\u0026rsquo; performance. Explore our detailed comic strip for a simplified and engaging explanation.\n","permalink":"https://blogsbykush.com/concept-breakdown/what-is-feature-engineering/","summary":"Discover the essentials of Feature Engineering in a classroom comic.","title":"What is Feature Engineering"},{"content":"My Reasons Psychology of Money by Morgan Housel is an international bestseller. I recently read this book after it was recommended by James Clear in his book \u0026ldquo;Atomic Habits.\u0026rdquo; Though I\u0026rsquo;ve always wanted to learn more about finance and money, I never got around to it. So, I picked up \u0026ldquo;Psychology of Money\u0026rdquo; to gain insights into investments and money management. However, this book offers much more than that; it teaches important life lessons intertwined with money matters using real-life examples. Learning from examples, in my view, is the most intuitive way to grasp concepts. Here, I\u0026rsquo;ll summarize what I\u0026rsquo;ve learned from the book, hoping it will be useful to you.\nThis book imparts important lessons: your habits, work ethic, routines, and attitude towards life and challenges ultimately determine your path and shape your identity.\nThe Lessons Learned The main idea of this book is that success with money has less to do with intelligence and more to do with behavior. Behavior is hard to teach, even to very smart people.\nFinancial success is not solely rooted in hard science; rather, it\u0026rsquo;s a soft skill where your behavior holds more significance than your knowledge.\nMorgan explains that when it comes to investment, no one is crazy. We all think we know how the world works, but we have only experienced a tiny part of it. Every financial decision makes sense to the person making it based on their own experiences.\nLuck and Risk One of my favorite chapters is \u0026ldquo;Luck and Risk.\u0026rdquo; Morgan suggests that every outcome in life is influenced by forces other than individual efforts. Luck and risk are similar, and you can\u0026rsquo;t believe in one without respecting the other.\nNothing is as good or as bad as it seems.\nMorgan argues that we often overlook the role of luck in success stories. Studying individual cases can be risky because extreme examples like billionaires or failures often don\u0026rsquo;t apply to most people. Instead, looking for broad patterns of success and failure provides more useful insights.\nLong-term Investing and Compounding The book emphasizes the importance of long-term investing and the power of compounding. Small, consistent actions over time can lead to significant wealth accumulation. Patience is key to financial success. Housel uses examples from nature to show how small changes can lead to substantial growth, highlighting the need to be vigilant and adaptable in financial endeavors.\nRetaining Wealth Acquiring money is one challenge, but retaining it is another, requiring patience and humility. To maintain wealth, adopting a survival mindset is crucial. This mindset prioritizes three key principles: patience, maintaining a margin of safety, and cultivating a balanced personality that is optimistic yet cautious.\nThe Importance of Tail Events Morgan devotes a chapter to the significance of tail events—extreme occurrences that happen rarely but have a huge impact. Success in investing depends on how you navigate these moments of fear and uncertainty.\nFreedom and Time Management Money can provide freedom, particularly in terms of time management. Financial resources allow you to have more control over your time, which is often linked to happiness. However, today\u0026rsquo;s generation struggles to disconnect from work, leading to a perceived loss of control over time despite being financially better off.\nThe Hidden Nature of Wealth Morgan clarifies that wealth often resides in financial assets that haven\u0026rsquo;t been converted into tangible possessions. People often want to spend a million dollars rather than save it, missing the essence of being wealthy.\nSaving and Humility Saving money is not tied to your income level but to your humility. Beyond a certain income level, what you truly need is often beneath your ego. Saving without a specific goal gives you options and flexibility, allowing you to wait for opportunities and seize them when they arise.\nBeing Reasonable The book suggests aiming to be reasonable rather than purely rational in financial decisions.\nOutliers and Surprises Things that have never happened before happen all the time.\nOutlier events often have the most significant impact. The most noteworthy events in history are the extreme outliers. Economic history is full of surprises, showing the unpredictability of the future. Recognizing that the future may differ significantly from the past is a valuable skill.\nRoom for Error The book provides examples to illustrate the concept of \u0026lsquo;Room for Error.\u0026rsquo; Recognizing uncertainty, randomness, and chance as inherent aspects of life is essential. A margin of safety acts as a buffer against unforeseen circumstances. While taking risks is necessary for progress, risking everything is never justifiable.\nThe most important part of every plan is planning on your plan not going according to plan.\nPoor Forecasters Morgan discusses how people are poor forecasters of their future selves. While imagining goals is easy, envisioning them amid real-life stresses is much harder. Since the future is uncertain, avoid extreme financial planning. Strive for balance to prevent future regret and promote endurance. Accepting that we will change our minds is crucial. The key is to embrace change and adapt quickly.\nIllusion of Control Everyone has an incomplete view of the world, yet we create complete narratives to fill in the gaps. The illusion of control is more persuasive than the reality of uncertainty, so we cling to stories that suggest outcomes are within our control. This focus on what we know, while neglecting what we don\u0026rsquo;t, makes us overly confident in our beliefs.\nLast Word \u0026ldquo;The Psychology of Money\u0026rdquo; offers several life lessons: seek humility when things go well and forgiveness and compassion when they don\u0026rsquo;t, as situations are never as good or as bad as they seem. Respect the power of luck and risk to focus on what you can control. Use money to gain control over your time. To improve as an investor, extend your time horizon—time is the most powerful force in investing. My favorite wisdom is that uncertainty, doubt, and regret are common costs in finance; we must define the cost of success and be ready to pay it. Additionally, embrace room for error, which may seem conservative but can keep you in the game and prove invaluable.\n","permalink":"https://blogsbykush.com/bookshelf/psychology-of-money-bookshelf/","summary":"Discover key lessons and insights from Morgan Housel\u0026rsquo;s \u0026lsquo;The Psychology of Money,\u0026rsquo; including the importance of behavior in financial success, the role of luck and risk, and the value of long-term investing.","title":"Psychology of Money - My Perspective and Key Takeaways"},{"content":"\nExplore the world of machine learning with our \u0026lsquo;Concept Breakdown\u0026rsquo; classroom series! In this conversation, a curious student dives into the realm of Gradient Descent, seeking to understand its role in evaluating model performance. The teacher paints a vivid analogy of navigating a foggy mountain, simplifying the intricate process of tweaking model parameters. Join us as we demystify complex concepts with humor and clarity\nDiscover the essence of Gradient Descent in machine learning through a conversation between a student and a teacher.\n","permalink":"https://blogsbykush.com/concept-breakdown/what-is-gradient-descent/","summary":"Discover the significance of Gradient Descent in machine learning through an engaging conversation between a student and teacher in our Concept Breakdown series.","title":"What is Gradient Descent"},{"content":"\nExplore the world of machine learning as we simplify the concept of hyperparameters in a classroom conversation. Understand their crucial role in optimizing model performance, with a relatable analogy using cooking instructions. Join us in demystifying complex topics in our Concept Breakdown series.\nDiscover the essence of hyperparameters in machine learning through a conversation between a student and a teacher.\n","permalink":"https://blogsbykush.com/concept-breakdown/what-is-hyperparameters/","summary":"Dive into the Concept Breakdown series conversation on hyperparameters in machine learning.","title":"What is Hyperparameters"},{"content":"Introduction Originally posted on istockphoto. Before you start reading this post, I\u0026rsquo;d like to offer a disclaimer: I\u0026rsquo;m not an expert in this field. I\u0026rsquo;m here to share my learning journey with the hope that it might help others embarking on this fascinating and sometimes daunting path.\nAbout five years ago, I set out to learn machine learning on my own. At the time, like many others, I was intrigued by buzzwords like Machine learning and Data Science. I wanted to explore something new, something beyond my daily work routine as a software engineer. However, I had no prior experience in the field of machine learning, which left me wondering if this path was right for me. To kickstart my journey, I enrolled in online courses, taking a few from platforms like Udemy and Coursera. These courses equipped me with the fundamentals, giving me the initial confidence that I could venture into this exciting domain.\nOnline Courses and Kaggle After traversing multiple online courses for a couple of months, I found myself growing weary. Many of these courses were lengthy and laden with theoretical content. This is when I stumbled upon Kaggle, and my journey took an exciting turn. I decided to give it a try and started with the very basic Housing Price competition. As I worked through the competition, I not only reviewed other participants\u0026rsquo; notebooks but also created my own. To my surprise, my submissions garnered praise and recognition. Although the problem statement was rather straightforward, it boosted my confidence significantly. I was able to apply theoretical knowledge, such as regression analysis, to real data, make predictions, and evaluate their accuracy in the context of predicting housing prices.\nAs I continued to tackle various problems and competitions, while also learning from the work of others, my skills improved. I developed a structured approach to solving machine learning problems, honed my analytical thinking, experimented with diverse strategies, and learned to adapt to evolving circumstances.\nIf anyone is interested in to kind of beginner\u0026rsquo;s competition I have participated in and the notebooks I have created, they can refer to my profile.\nReal-World Projects After a few months on Kaggle, I began searching for real-world use cases that organizations and companies were trying to address using machine learning. I approached my manager with this idea, and after an initial discussion, he provided me with several problem statements to investigate. This was during a time when organizations were keen on harnessing machine learning to enhance their existing workflows. I decided to work on a Regression Test Case Selection by Machine Learning proof of concept (POC). This experience taught me how to handle data effectively and how a machine-learning solution could translate into a practical business solution. It also bolstered my credibility in handling machine learning problems from the ground up. Engaging with real-world use cases can significantly enhance your learning journey. However, compelling projects may not always be readily available. In such instances, I stumbled upon the Project Pro Repository, which I found to be highly efficient. It offers a wide range of projects suitable for beginners to intermediates.\nCertificates and Blogging After the successful completion of the POC, I found myself taking on side projects and other POCs at my workplace. These experiences were invaluable for gaining a better understanding of data handling and how to approach various problem statements. However, my learning remained somewhat unstructured. I knew many concepts but lacked a cohesive framework. At this point, while working with AWS cloud services, I decided to pursue the AWS Machine Learning Specialist Cloud Certification. Preparing for this certification not only introduced me to new concepts in a structured manner but also enhanced my understanding of machine learning-related cloud services. In my current role, where I routinely deploy machine learning models and pipelines in the AWS cloud, this certification has served as the fundamental building block for my work.\nIf anyone is interested in preparing for this certification, they can refer to this blog for a detailed summary of my preparation.\nIn addition to all this, I made a conscious effort to document my work and share my experiences through blogging. Writing about your work can be challenging, and in my opinion, if you can explain a concept most thoroughly, it signifies a deep understanding. Initially, I contributed to various publications on Medium, but I eventually launched my blog site, blogsbykush to maintain accountability and share my learning in the most straightforward manner possible.\nLast Word As you can see, I\u0026rsquo;ve tried a multitude of approaches, and I\u0026rsquo;m continually exploring new avenues for learning. While the learning journey may vary from one individual to another, experimentation remains a key aspect of mastering any new field. Engaging in projects, POCs, and hackathons is invaluable compared to relying solely on theoretical knowledge and online courses. Resources like Kaggle, official POCs, project repositories (such as the one available here), and many others can provide tremendous support on your learning journey.\nIn conclusion, while I may not be an expert, my journey through machine learning has been rewarding. It\u0026rsquo;s been filled with experimentation, hands-on experiences, structured learning, and the joy of sharing knowledge. I hope my story encourages you to take the plunge into this dynamic field and embrace the most effective way to learn machine learning - through practice and application.\n","permalink":"https://blogsbykush.com/what-is-the-most-effective-way-to-learn-ml/","summary":"Learn about the journey of a beginner in the field of machine learning. Discover how they started with online courses, transitioned to Kaggle competitions, tackled real-world projects, pursued certifications, and documented their experiences through blogging. This article provides insights and encouragement for those looking to embark on a similar learning journey.","title":"What is the most effective way to learn Machine Learning"},{"content":"\nDiscover the essence of cross-validation in machine learning through a conversation between a student and a teacher.\n","permalink":"https://blogsbykush.com/concept-breakdown/what-is-cross-validation/","summary":"Learn about cross-validation in machine learning through a fun and relatable analogy.","title":"What is Cross Validation"},{"content":"\nExplore Computer Vision in our Concept Breakdown series as we break down this complex topic into simple terms. Learn how it empowers computers to \u0026lsquo;see\u0026rsquo; and understand the world, just like teaching a child to recognize objects. Discover applications like Facial Recognition and Self-Driving Cars.\nExplore Computer Vision in our Concept Breakdown series as we break down this complex topic into simple terms. Learn how it empowers computers to \u0026lsquo;see\u0026rsquo; and understand the world, just like teaching a child to recognize objects.\n","permalink":"https://blogsbykush.com/concept-breakdown/what-is-cv/","summary":"Demystify Computer Vision with our Concept Breakdown series.","title":"What is Computer Vision"},{"content":"\nNatural Language Processing (NLP) is like teaching computers to understand and talk like humans, enabling them to comprehend our words, discern intentions, and respond meaningfully. It\u0026rsquo;s akin to preschool teachers communicating with children from different languages.\n","permalink":"https://blogsbykush.com/concept-breakdown/what-is-nlp/","summary":"Learn about Natural Language Processing (NLP) in simple terms.","title":"What is Natural Language Processing"},{"content":"\nA conversation explaining recommendation system, its definition, examples, and a relatable analogy for better understanding.\nThe conversation introduces recommender systems as technology that uses machine learning for personalized suggestions, exemplified by Netflix\u0026rsquo;s content recommendations. An analogy likens it to a thoughtful grandmother\u0026rsquo;s choices, highlighting the role of machine learning in tailored recommendations.\n","permalink":"https://blogsbykush.com/concept-breakdown/what-is-recommender-system/","summary":"Explore the concept of recommender systems through a student-teacher dialogue.","title":"What is Recommender system"},{"content":" A conversation explaining transfer learning, its definition, examples, and a relatable analogy for better understanding.\nA conversation explaining transfer learning, a technique in machine learning that reuses knowledge from one problem to improve performance on related problems. Explore real-world examples and a relatable analogy to understand the concept easily\n","permalink":"https://blogsbykush.com/concept-breakdown/what-is-transferlearning/","summary":"Learn about transfer learning, a technique in machine learning that reuses knowledge from one problem to improve performance on related problems.","title":"What is Transfer Learning"},{"content":" In this conversation, we explore about deep learning and neural networks. Through a conversation between a teacher and a student, we uncover the essence of deep learning and its advanced capabilities. Discover how deep learning uses deep neural networks to automatically learn complex patterns from data, and how neural networks, inspired by the human brain, process information and make predictions. Dive into the analogy of working in a team and the game of \u0026lsquo;Chinese Whispers\u0026rsquo; to gain a deeper understanding of these concepts\n","permalink":"https://blogsbykush.com/concept-breakdown/what-is-deeplearning-and-neuralnetwork/","summary":"Uncover the power of deep learning and neural networks in this engaging conversation between a teacher and a student.","title":"What is Deep Learning and Neural Network"},{"content":" In this conversation, the teacher and student discuss different types of Machine Learning. The student explains supervised learning, unsupervised learning, and reinforcement learning, providing simple examples and analogies to illustrate each type. The teacher acknowledges the student\u0026rsquo;s clear and concise explanations.\n","permalink":"https://blogsbykush.com/concept-breakdown/type-of-machine-learning/","summary":"Dive into the conversation about different types of Machine Learning, including supervised learning, unsupervised learning, and reinforcement learning.","title":"Types of Machine Learning"},{"content":"In this post, I explore AWS Sagemaker\u0026rsquo;s features and use cases in Machine Learning. I am focusing on its capabilities that enable early deployment of Machine Learning models, significantly reducing time to production.\nIntroduction Amazon SageMaker is a fully managed machine learning service that empowers data scientists and developers to swiftly build, train, and deploy machine learning models. With SageMaker, the process of building and training models becomes effortless as it eliminates the need for server setup or infrastructure management. The service provides an easy-to-use Jupyter notebook for seamless data exploration and analysis, capable of efficiently handling massive datasets.\nWhat sets SageMaker apart is its flexibility, supporting custom algorithms and frameworks that can adapt to specific preferences and goals. Moreover, it offers two distinct interfaces: SageMaker Studio and SageMaker Console, which facilitate the management and execution of machine learning workflows\nUse Case One of the common challenges in the machine-learning world is\nHow to reduce the time to production for ML-driven solutions\nA typical machine learning project begins with a Proof of Concept (POC) for a business use case. Usually, the POC is developed in notebooks like Jupyter or Colab, either locally or in private cloud accounts. Once promising outcomes are achieved in the POC, the team decides to transition the solution into production. However, this conversion phase from POC to Production often involves reworks in areas such as Data Engineering, Data Pipelines, and Model Evaluation, which inevitably increase the time required to reach production.\nAccording to a study by Alegion and Dimensional Research, one-third of AI/ML projects encounter roadblocks and stall at the proof of concept phase.\nThe following graphics support the article’s analysis:\nResearch conducted by Alegion and Dimensional and published in The News Stack. We can overcome these challenges by embracing the early adoption of AWS SageMaker.One of the standout features of SageMaker is its ability to streamline the deployment of models into a production-ready environment. By leveraging this comprehensive end-to-end machine learning cloud solution, Engineers and Teams can seamlessly experiment, build, train, and deploy ML models within a unified environment.\nTo illustrate this, I have created a Jupyter notebook where I have used SageMaker Studio to build, train, deploy, and monitor an XGBoost model. I have covered the entire Machine Learning workflow from feature engineering and model training to batch and live deployments for ML models. By utilizing SageMaker\u0026rsquo;s capabilities, I have tried to highlight how teams can efficiently navigate through the different stages of model development and deployment, accelerating the time to market.\nImplementation In the notebook you will discover how a single notebook instance allows us to accomplish the following tasks seamlessly :\nDownloading and processing the dataset stored in remote cloud storage (S3 bucket) Creating a SageMaker experiment to train the model Evaluating the model\u0026rsquo;s performance Deploying the model as RESTful HTTPS endpoints Monitoring the model In SageMaker Studio notebooks are one-click Jupyter notebooks that contain everything you need to build and test your training scripts.\nI have created a SageMaker Notebook, downloaded the dataset, and then upload the dataset to Amazon S3 bucket.After downloading and staging the dataset in S3 bucket, I created the SageMaker Experiment to train the model.\nAfter training the model, I created offline/batch inference to evaluate model performance on unseen data. Evaluation metrics helped me to tune the model for better results.\nWith the help of SageMaker SDK libraries, I was able to deploy the best-performing model as a RESTful HTTP endpoint and monitor the deployed endpoints for any data drift.\nAll of this happened in a single notebook with the bare minimum code.\nFeatures As an end-to-end machine learning cloud solution, SageMaker comes with a plethora of features. However, here are a few standout features that are worth mentioning:\nInfrastructure/Resource Management Amazon SageMaker handles the management of an S3 bucket for storage and separate compute instances for processing. Processing tasks run on dedicated compute instances that are independent of the notebook instance. This allows for seamless experimentation and code execution in the notebook while processing is in progress. Upon completion of the processing job, SageMaker automatically terminates the instances, ensuring efficient resource utilization.\nSageMaker\u0026rsquo;s ephemeral cluster approach means that each training job has its own dedicated EC2 instances, which are active for the duration of the training process. Development and training occur in separate EC2 instances, avoiding conflicts with virtual environments, package installations, and resource dependencies.\nSageMaker also provides a notebook instance, which is an ML compute instance running the Jupyter notebook app. SageMaker handles the creation of instances and related resources, allowing you to use the notebook for data preparation and processing, code development, model training, deployment, testing, and validation.\nExperiment Management A SageMaker Experiment is a collection of processing and training jobs related to the same machine-learning project. It is a capability that lets you organize, track, compare, and evaluate your machine-learning experiments.\nEach training job is logged as a trial within SageMaker, representing an iteration of the end-to-end training process. This includes pre-processing and post-processing tasks, datasets, and relevant metadata. With SageMaker\u0026rsquo;s experiment management capabilities, multiple trials can be included in a single experiment, making it easy to track and compare different iterations over time.\nSageMaker Python SDK SageMaker provides a Python library that simplifies model training and deployment. The library abstracts platform specifics, offering convenient methods and default parameters like deploy() and fit() for model deployment and training.\nSample Notebooks SageMaker provides various Jupyter notebooks that demonstrate training and deployment of models using specific algorithms and datasets. You can start with a suitable notebook and modify it to accommodate your own data source and specific requirements.\nChoices to train model SageMaker offers flexibility in training your machine learning models. Here are a few options:\nScript Mode: SageMaker manages open-source containers such as mxnet, TensorFlow, PyTorch, scikit-learn, and more. Docker Containers: Bring Your Own (BYO) Model by using your own Docker file and managing user-specific Docker images and registry within ECR. AWS ML Marketplace: Access different algorithms and pre-trained model artifacts through a subscription model in the free tier, allowing for training your own data or using pretrained models. Built-in Algorithms: SageMaker provides a suite of algorithms, pre-trained models, and pre-built solution templates to quickly get started with training and deploying machine learning models. Cost and Learning Curve Amazon SageMaker offers a flexible and transparent pricing model that enables you to pay only for the resources you use. Data storage costs are based on the amount of data stored in Amazon S3 or other supported storage options. When it comes to training instances, you are charged for the compute resources utilized during model training.\nAWS SageMaker provides a range of instance types with varying capabilities, and the pricing is dependent on factors such as instance type, duration of usage, and the number of instances employed.\nSimilarly, for model deployment, you are billed for the resources utilized to host your models in a production environment.\nIn my demo implementation, I utilized the ml.t3.medium notebook instance, which cost approximately 1.2 USD per day. While I paid a minimal amount for a demo or one-time activity, it is important to be mindful of SageMaker billing. In my experience, without proper cost optimization procedures, all cloud services can potentially result in high bills.\nTo effectively manage and optimize your costs with SageMaker, I recommend leveraging cost optimization features, including automatic model scaling and instance selection recommendations.\nWhen it comes to learning, AWS offers extensive documentation, tutorials, and sample notebooks to facilitate the learning process. Furthermore, the availability of pre-built algorithms and optimized frameworks simplifies the implementation of machine learning models.\nWith its intuitive features, extensive libraries, and extensive community support, SageMaker helps reduce the learning curve. In many scenarios, ML Engineers don\u0026rsquo;t require advanced machine learning skills to access and leverage various ML functionalities in SageMaker\nLast Word AWS SageMaker offers a powerful end-to-end machine learning cloud solution, allowing engineers to seamlessly build, train, deploy, and monitor machine learning models within a single platform. By incorporating proper cost optimization techniques, SageMaker can prove to be a cost-efficient solution for organizations. While this post covers only a few standout features, there are many other capabilities to explore.\nFor further information, I recommend referring to the AWS Official Documentation for SageMaker and the Official SageMaker Deep Dive Tutorials Series by Emily Webber.\nIn competition with AWS SageMaker, Google and Microsoft also provide their own end-to-end machine learning cloud solutions. Google Cloud Datalab is a standalone serverless platform for building and training machine learning models. Meanwhile, Microsoft Azure Machine Learning Studio presents tough competition to SageMaker as Microsoft expands its services. Comparative studies suggest that SageMaker is the better option if you are well-versed in programming, especially for complex and large-scale projects. On the other hand, Azure Machine Learning Studio is more suitable for those with smaller and simpler goals.\n","permalink":"https://blogsbykush.com/mlops/sagemaker-empowering-your-ml-lifecycle/","summary":"In this blog post, I explore the features and use cases of AWS SageMaker in the field of Machine Learning. Specifically, I focus on its capabilities that enable early deployment of ML models, significantly reducing the time to production. Discover how SageMaker simplifies the process of building, training, and deploying models, and learn about its standout features, such as infrastructure/resource management, experiment management, the SageMaker Python SDK, and more. Find out about the cost and learning curve associated with SageMaker, and gain insights into its competition with Google Cloud Datalab and Microsoft Azure Machine Learning Studio","title":"AWS SageMaker - Empowering Your ML Lifecycle"},{"content":"\nIn this conversation, a teacher and student discuss the concept of Machine Learning. The student explains that Machine Learning involves teaching computers to learn and make decisions from data without explicit programming. They provide real-life examples, such as personalized recommendations on platforms like Netflix and Spotify. The conversation concludes with a simplified explanation comparing Machine Learning to how our brains learn from experiences\n","permalink":"https://blogsbykush.com/concept-breakdown/what-is-machine-learning/","summary":"Explore a conversation between a teacher and student as they delve into the world of Machine Learning.","title":"What is Machine Learning"},{"content":" In this classroom discussion, the teacher introduces the topic of evaluation metrics, and a student provides an explanation. They discuss different evaluation metrics used in machine learning, such as accuracy, precision, and recall. The student then simplifies the concept using an analogy of sorting crayons into different piles with the help of a friend, relating accuracy, precision, and recall to the friend\u0026rsquo;s performance in the task\n","permalink":"https://blogsbykush.com/concept-breakdown/evaluation-metrics/","summary":"Understand the concept of evaluation metrics in machine learning through this classroom conversation.","title":"Evaluation Metrics"},{"content":"\nIn this classroom discussion, the teacher asks the students to explain data leakage and its challenges. A student describes data leakage as the inclusion of information in the training data that would not be available in the real world. The student further explains that data leakage is challenging because it is often difficult to identify and can result in models that perform poorly in real-life scenarios. The student then provides a simple analogy of studying for a test based on specific information, only to realize that the actual exam doesn\u0026rsquo;t cover that information. The teacher acknowledges the analogy and relates it to the concept of overfitting and poor performance on unseen data\n","permalink":"https://blogsbykush.com/concept-breakdown/data-leakage/","summary":"Understand the concept of data leakage and its challenges through this classroom conversation.","title":"Data Leakage"},{"content":"\nIn this classroom discussion, the teacher introduces the concept of class imbalance, where one class or label has more or fewer examples than another. A student explains that class imbalance can make it challenging for machine learning models to make accurate predictions. The student suggests two methods to deal with class imbalance: undersampling, which involves removing examples from the majority class, and oversampling, which involves creating new examples of the minority class. To make it simpler, the student uses the analogy of making teams with friends, where balancing the number of players on each team can be achieved by either removing a friend from one team (undersampling) or adding a friend to the other team (oversampling). The teacher acknowledges the student\u0026rsquo;s explanation and expresses satisfaction with the analogy\n","permalink":"https://blogsbykush.com/concept-breakdown/class-imbalance/","summary":"Understand the concept of class imbalance in machine learning and learn how to address it effectively.","title":"Class Imbalance"},{"content":"My Reasons Atomic Habits by James Clear is a New York Times bestsellers. This famous book caught my attention during the pandemic. It was becoming increasingly difficult for me to balance my professional and personal work during this time. While I am not a fan of \u0026lsquo;self-help\u0026rsquo; books, I really want to improve my daily routine, so I want to give this book a try. Honestly, I am still far from having an ideal routine, but after reading this book, I know my strengths and areas for improvement, and I have clarity about where I need to improve. Additionally, I was caught by the tagline\nTiny changes, remarkable results\nHow can small changes in our routine impact our long-term goals? In order to understand these things, I chose the book.\nThe Lesson Learned \u0026ldquo;Atomic Habits\u0026rdquo; is a practical guide on how to form and maintain good habits and break bad ones. The author draws on the latest research in psychology, neuroscience, and behavioral economics to explain why habits are so powerful and how we can use this knowledge to make positive changes in our lives.\nIt begins by examining why tiny changes make a big difference, the intent behind the name - \u0026lsquo;Atomic Habits\u0026rsquo; is :\nHabits are like the atoms of our lives.Each one is a fundamental unit that contributes to your overall improvement\nThe author discussed how incremental changes are often underestimated when we focus on big goals. Examples were provided that illustrated how habits can compound and work against you.\nThe author argues that our habits are often a reflection of our identity. In other words, if we want to change our habits, we need to change our sense of self. For example, instead of saying “I want to lose weight,” say “I am the type of person who makes healthy choices.”\nOne of the interesting things I have learned is The science of habit formation, it can be broken down into four stages, known as the habit loop, it consists:\nCue: This is the trigger that prompts the habit. Cues can be anything that your brain associates with a particular behavior or activity. For example, seeing a donut might be a cue for you to eat it, or hearing your phone buzz might be a cue for you to check your messages. Craving: They are the motivational force behind every habit. Without some level of motivation or desire, we have no reason to act. What you crave is not the habit itself but the change in state it delivers. For example, you do not crave smoking a cigarette, you crave the feeling of relief it provides. Response: This is the actual habit you perform which can take the form of a thought, or an action. Reward: This is the positive outcome that you get from engaging in the habit. Rewards can be physical, emotional, or psychological. For example, the reward for eating a donut might be the taste and feeling of satisfaction that comes from eating something sweet and indulgent. If the behavior is insufficient in any of the four stages, it will not become a habit. Without the first three, a behavior will not occur and without all four behavior will not be repeated. In summary :\ncue triggers a craving which motivates a response, which provides a reward\nThis four-step process is not something that happens occasionally, but rather it is an endless feedback loop that is running and active during every moment you are alive.\nThis habit loop laid the foundation of Four laws of behavior change,which are:\nMake it obvious, attractive, easy, and satisfying\nMake it Obvious : The author suggests that to develop a new habit, it\u0026rsquo;s important to make it obvious and visible. This can be done by using cues, such as placing your gym clothes next to your bed or setting an alarm to remind you to take a break from work. Making your habits more visible can help your brain recognize and remember them, making it easier to stick to them over time.\nMake it attractive : To develop a new habit, make it more enjoyable by linking it to something you already like. For instance, exercise while listening to music if you enjoy it. Or, join a group or find an accountability partner to make it more social. By making your habits more enjoyable, you increase the likelihood of sticking to them long-term.\nMake it Easy : The author suggests that to develop a new habit, helps to make it easy to do. This can be done by breaking it down into smaller, manageable steps and reducing the barriers to starting. To make a habit easier, use the Two-Minute Rule. This means making the first two minutes of a new habit as easy as possible. For example, if you want to start meditating daily, start with two-minute meditations instead of longer ones. This will make it easier to get started. Overall, by making a habit easier to start and maintain, you increase the likelihood that you\u0026rsquo;ll stick with it over time and make it a part of your daily routine.\nMake it satisfying : Making a habit rewarding and satisfying is essential to maintaining it in the long run. Keeping track of your progress and celebrating your successes along the way can help you achieve this. You can also reward yourself for completing a task right away, such as with a small treat or a break afterward. It becomes easier to maintain habits over time when you reinforce the positive feelings associated with them.\nLast Word \u0026ldquo;Atomic Habits\u0026rdquo; is an essential read for those seeking lasting, positive change in their lives. James Clear\u0026rsquo;s evidence-based approach to habit formation breaks the process down into easy-to-understand steps. His engaging writing style makes it enjoyable to read. \u0026ldquo;Atomic Habits\u0026rdquo; provides actionable advice and insights that will help improve your health, productivity, or relationships, and achieve your goals to build a better life.\n","permalink":"https://blogsbykush.com/bookshelf/atomic-habits-bookshelf/","summary":"Discover the valuable lessons learned from reading Atomic Habits by James Clear. Explore the author\u0026rsquo;s evidence-based approach to habit formation and how small changes can lead to remarkable results. Learn about the science of habit formation, the habit loop, and the four laws of behavior change. Find out how to make habits obvious, attractive, easy, and satisfying, and develop a better daily routine. Improve your health, productivity, and relationships with actionable advice from this essential book","title":"Atomic Habits - My Perspective and Key Takeaways from James Clear's Book"},{"content":"Privacy Policy The privacy of my visitors is extremely important. This Privacy Policy outlines the types of personal information that is received and collected and how it is used.\nFirst and foremost, I will never share your email address or any other personal information to anyone without your direct consent.\nLog Files Like many other websites, this site uses log files to help learn about when, from where, and how often traffic flows to this site. The information in these log files include:\nInternet Protocol addresses (IP) Types of browser Internet Service Provider (ISP) Date and time stamp Referring and exit pages Number of clicks All of this information is not linked to anything that is personally identifiable.\nCookies and Web Beacons When you visit this site \u0026ldquo;convenience\u0026rdquo; cookies are stored on your computer when you submit a comment to help you log in faster to Disqus the next time you leave a comment.\nThird-party advertisers may also place and read cookies on your browser and/or use web beacons to collect information. This site has no access or control over these cookies. You should review the respective privacy policies on any and all third-party ad servers for more information regarding their practices and how to opt-out.\nIf you wish to disable cookies, you may do so through your web browser options. Instructions for doing so can be found on the specific web browsers\u0026rsquo; websites.\nGoogle Analytics Google Analytics is a web analytics tool I use to help understand how visitors engage with this website. It reports website trends using cookies and web beacons without identifying individual visitors. You can read Google Analytics Privacy Policy.\nGoogle Adsense Google Adsense, a third party affiliate marketing network, uses cookies to help make sure I get a commission when you buy a product after clicking on a link or ad banner that takes you to the site of one of their merchants. You can read Google Adsense Privacy Policy.\nDisclosure Policy I make money on this website through affiliate programs. If you click an affiliate link or ad banner and buy the product, you help support this website because I\u0026rsquo;ll get a percentage of that sale.\nCurrently I\u0026rsquo;m an affiliate for Amazon and Google Adsense.\nWhat this means for you:\nI became an affiliate to earn revenue towards the costs of running and maintaining this website. Where I have direct control over which ads are served on this website I offer only products that are directly related to the topic of this website and products that a reader/subscriber would have a genuine interest in or need of. I do not and will not recommend a product just for the sake of making money. I do not let the compensation I receive influence the content, topics, posts, or opinions expressed on this website. I respect and value my readers too much to write anything other than my own genuine and objective opinions and advice. Just like this website, my Disclosure Policy is a work in progress. As the revenue streams evolve, so will this page.\n","permalink":"https://blogsbykush.com/terms/","summary":"\u003ch2 id=\"privacy-policy\"\u003ePrivacy Policy\u003c/h2\u003e\n\u003cp\u003eThe privacy of my visitors is extremely important. This Privacy Policy outlines the types of personal information that is received and collected and how it is used.\u003c/p\u003e\n\u003cp\u003eFirst and foremost, I will never share your email address or any other personal information to anyone without your direct consent.\u003c/p\u003e\n\u003ch3 id=\"log-files\"\u003eLog Files\u003c/h3\u003e\n\u003cp\u003eLike many other websites, this site uses log files to help learn about when, from where, and how often traffic flows to this site. The information in these log files include:\u003c/p\u003e","title":"Terms and Privacy Policy"},{"content":"\nIn this classroom discussion, the teacher introduces the concepts of overfitting and underfitting in machine learning models. A student explains that overfitting occurs when the model is too complex and fits the training data too closely, resulting in poor generalization to new data. The student uses the analogy of a familiar route to work, where an unexpected traffic jam causes a delay, highlighting the over-reliance on past experiences. The teacher acknowledges the student\u0026rsquo;s explanation. The student then explains underfitting as a situation where the model is too simple and fails to capture important patterns in the data, leading to poor performance. The student uses the analogy of taking a multiple-choice exam without preparation, where guessing without any information results in mostly incorrect answers. The teacher commends the class for their understanding of the concepts\n","permalink":"https://blogsbykush.com/concept-breakdown/overfitting-and-underfitting/","summary":"Understand the concepts of overfitting and underfitting in machine learning models.","title":"Overfitting Vs Underfitting"},{"content":"\nIn this classroom discussion, the teacher asks the students to explain what a model is in the context of machine learning. One student describes a model as an abstract representation of a process that uses known information to predict responses in different situations. Another student provides a simpler explanation, comparing a model to a robot that learns from examples to recognize objects. The teacher acknowledges the students\u0026rsquo; explanations and adds that our minds also rely on mental models to predict real-time situations and guide decision-making\n","permalink":"https://blogsbykush.com/concept-breakdown/what-is-model/","summary":"Understand the concept of a model in machine learning through a classroom discussion.","title":"Models - What are they"},{"content":"My Objective and Preparation Strategy to clear this Specialty Exam\nIntroduction Last week I had cleared AWS Certified Machine Learning Specialty Exam and after that, I am getting a lot of questions about exams and how did I prepare for this exam. In this post, I am sharing my objective of taking this exam, preparation strategy, and resources I have referred. I hope it will be helpful to someone who is planning to take this certification.\nObjective for taking the exam I am working in the Machine Learning domain for almost two years and most of my learning came either from online courses or Pet Projects/POCs(Proof Of Concepts), which is not a structured way of learning so as a result I was left with many weak concepts and understanding gap. The curriculum for this certification brings Structure Learning as it covers the whole Machine Learning Cycle along with AWS-specific services.\nAnother reason behind taking this exam is to gain knowledge for building and deploying End to End Machine Learning Products in Cloud. The majority of my work is POCs and Pet projects which most of the time ends up in Juypter notebook in GIT and never deployed to Production, this preparation helps me to bridge the gap of building reliable and effective machine learning architecture in the cloud.\nTip: If you want to focus more on the ML domain then start your preparation with that and then move to AWS specific content, this will help you to build confidence in the domain\nAbout the Exam As this is a Specialist exam and emphasis a lot on the domain, this is different from other AWS certifications and needs proper planning and time allocation. The exam is split into 4 domains: Data Engineering (20%), Exploratory Data Analysis(EDA) (24%), Modelling (36%), and ML Implementation \u0026amp; Operations (20%). The content within EDA and Modelling is a balance of domain and specific AWS services, whereas the Data Engineering is more on the AWS side. For MLOps, I have found that the majority of the content got covered in the Modelling section. They are expecting real-time experience in the Machine learning field so make sure that you have some working experience in the field.\nTip : I have booked the exam date in advance before starting my preparation, this helped me to have time-bound and focused preparation. You will have two chances to postpone the exam, but try to avoid rescheduling\nPreparation Strategy and Resources I have put in close to 185 hours in the span of 4 months to prepare for this exam, referred following online material in the same order\nAWS Resources — I have started from Machine Learning Path that AWS suggests taking to prepare for the exam. From this path, I have picked only The Elements of Data Science, and the Exam Readiness course. The first one is a very basic and refresher course and the second one talks about the content that is expected to know in each of the domains and their sub-domains, use this as a reference to check your readiness for the exam Udemy — AWS Certified Machine Learning Specialty 2021 — Hands On! By Stephane Maarek and Frank Kane. This is a deep dive course and gives you the full picture of exam content Sage Maker — AWS Official Documentation and Amazon SageMaker Technical Deep Dive Series . Amazon SageMaker is a fully managed machine learning service and very much important in terms of exam content, Modelling and ML Ops sections are heavily dependent on this service. It’s a common piece of advice from another exam taker to read plenty of SageMaker official documentation, but I found this one is a little intimating as it is very lengthy. To understand the service thoroughly I have referred to deep-dive videos which are series of 16 videos and I found them very useful. Along with the above content, I have referred to some generic available articles/blog posts to understand the concept and prepared notes in this Trello board. These notes are distributed in four Lists, each list represents one domain of the exam. Each note is labelled with priority along with resources links, screenshots. This board becomes a Kanban board which helped me to stay on top of my learning.\nTrello Board Sample Card Tip: As there is so much to learn, it is easy to get overwhelmed and lose track of what you have already learned and what needs to be covered next. In this case, your notes will come to the rescue. I will recommend preparing notes for this exam, it will also help you in last-minute revision\nPractice Exam In last week of my preparation I have given around 10 Practice set which helped me to understand my weak topics and where I need to focus more. I have put links along with recommendation for all these mock/practice test in ‘Resource’ list of Trello board .\nReal Exam I was a little skeptical before the actual exam date as I have not secured a great score in most of the practice exams and I am lacking in confidence. Since I had committed to myself not to change the exam date, so I went for it and passed the exam with 86.5%. During the exam, I have used the Process of Elimination extensively to remove wrong choices and I guess that worked for me because in the mock exam I was not using this technique.\nTip: Take your mock/practice exam seriously as it will help you to build your confidence for the real exam. Try to give mock exams in real-time exam like time-bound manner\nOther resources Apart from above-mentioned resources, I have referred following blogs which helped me to get started and sources to useful materials\nBlog by Raghav Dave Similar blog post by Simple Data flow I have also referred book Hands-on Machine Learning with Scikit-Learn, Keras, and TensorFlow to understand some core concepts Last word Along with lengthy and difficult certification to achieve this one is expensive also, so be clear in your objective —why you want to take this exam. As mentioned, my objective was to have structured learning to fill the gap in my knowledge and understanding to build reliable and effective machine learning architecture in the cloud, make sure you have your valid reason and objective before committing for this specialist certification.\nDisclaimer: Just to let you know, this blog post was originally published on Medium. If you\u0026rsquo;d like to check out the original, you can find it at this link.\n","permalink":"https://blogsbykush.com/how-i-prepared-for-aws-ml-specialty-certification/","summary":"Discover my objective and preparation strategy for clearing the AWS Certified Machine Learning Specialty Exam. Learn about the structure of the exam, its domains, and the importance of real-time experience in the machine learning field. Find out the online resources and courses I referred to during my preparation, including AWS resources, Udemy courses, SageMaker documentation, and additional articles. Explore the practice exams I took and how they helped me identify my weak areas. Gain insights into the tips and techniques I used during the real exam to pass with a high score. Get advice on preparing notes and staying organized throughout the learning process","title":"How I Prepared for AWS ML Specialty Certification"},{"content":"Introduction As COVID-19 is wreaking havoc, every Data Scientist is trying to get some insights from related data. In this post, I am also attempting to get answers to the following questions with the help of Data Visualization.\nWhat we can learn from countries like China, South Korea, and Italy the most impacted countries\nWhat is the current state in countries where it started late, like India and US, and what could be the possible future state in these two countries\nThe intent is to get basic insights without complex graphs and code and to keep things simple and easy to understand I have used only Plotly Scatter and Bar graphs.\nThis is the summary of my Kaggle notebook in which I have explored the COVID-19 data set provided by John Hopkins University. In case interested in more detailed code and explanation please refer to this notebook, where data is getting updated on daily basis.\nHow China, Italy, and South Korea are impacted Let’s start analyzing the situation in China first\n#import main data set df=pd.read_csv(\u0026#39;../input/novel-corona-virus-2019-dataset/covid_19_data.csv\u0026#39;) print(\u0026#34;Importing the data set and verifying the columns\u0026#34;) #Converting date column into correct format df[\u0026#39;ObservationDate\u0026#39;]=pd.to_datetime(df[\u0026#39;ObservationDate\u0026#39;]) #Geting latest timestamp latest_timestamp=df.iloc[df.last_valid_index()][\u0026#39;ObservationDate\u0026#39;] #Geting latest date latest_date=latest_timestamp.strftime(\u0026#34;%Y-%m-%d\u0026#34;) print(\u0026#34;Data available till {}\u0026#34;.format(latest_date)) #Checking Columns in data set df.head() Let’s add one more column here, Active Cases (Confirmed Cases-Recovered-Deaths)\n#Adding Active cases df[\u0026#39;ActiveCases\u0026#39;] = df[\u0026#39;Confirmed\u0026#39;] - df[\u0026#39;Deaths\u0026#39;] - df[\u0026#39;Recovered\u0026#39;] print(\u0026#34;Active Cases Column Added Successfully\u0026#34;) df.head() As data is present across provinces in China, we will group the data by Observation Date and add one more column ‘Week’ to see the trend across the weeks\n#Taking out of data for China region df_china=df[df[\u0026#39;Country/Region\u0026#39;] == \u0026#39;Mainland China\u0026#39;] print(\u0026#34;Number of Records for China are {}\u0026#34;.format(df_china.shape)) #Group by date df_china=df_china.groupby([\u0026#39;ObservationDate\u0026#39;]).sum().reset_index() print(\u0026#34;Grouped by Observation Date\u0026#34;) df_china[\u0026#39;Week\u0026#39;] = df_china[\u0026#39;ObservationDate\u0026#39;].dt.week print(\u0026#34;Week Column Added successfully\u0026#34;) df_china.head() Let’s visualize the trend and latest status in China\n#Visualizing the trends fig = go.Figure() fig.add_trace(go.Scatter(x=df_china[\u0026#39;ObservationDate\u0026#39;], y=df_china[\u0026#39;Confirmed\u0026#39;], mode=\u0026#39;lines\u0026#39;,name=\u0026#39; Confirmed Cases\u0026#39;)) fig.add_trace(go.Scatter(x=df_china[\u0026#39;ObservationDate\u0026#39;], y=df_china[\u0026#39;Deaths\u0026#39;], mode=\u0026#39;lines\u0026#39;,name=\u0026#39;Deaths\u0026#39;)) fig.add_trace(go.Scatter(x=df_china[\u0026#39;ObservationDate\u0026#39;], y=df_china[\u0026#39;Recovered\u0026#39;], mode=\u0026#39;lines\u0026#39;,name=\u0026#39;Recovered Cases\u0026#39;)) fig.add_trace(go.Scatter(x=df_china[\u0026#39;ObservationDate\u0026#39;], y=df_china[\u0026#39;ActiveCases\u0026#39;], mode=\u0026#39;lines\u0026#39;,name=\u0026#39;Active Cases\u0026#39;)) fig.update_layout(title_text=\u0026#39;Trend in China\u0026#39;,plot_bgcolor=\u0026#39;rgb(250, 242, 242)\u0026#39;) fig.show() #Getting the latest observation date from data set df.iloc[df.last_valid_index()][\u0026#39;ObservationDate\u0026#39;] #Taking out record for latest records in separate data frame df_china_lately=df_china[df_china[\u0026#39;ObservationDate\u0026#39;] == latest_date] #Creating temporary data frame for visualizing purpose df_china_status=pd.DataFrame(columns=[\u0026#39;Numbers\u0026#39;,\u0026#39;Status\u0026#39;]) #Appending Records from latest data frame df_china_status=df_china_status.append({\u0026#39;Numbers\u0026#39;:df_china_lately.iloc[0][\u0026#39;Confirmed\u0026#39;],\u0026#39;Status\u0026#39;:\u0026#39;Confirmed\u0026#39;},ignore_index=True) df_china_status=df_china_status.append({\u0026#39;Numbers\u0026#39;:df_china_lately.iloc[0][\u0026#39;ActiveCases\u0026#39;],\u0026#39;Status\u0026#39;:\u0026#39;Active Cases\u0026#39;},ignore_index=True) df_china_status=df_china_status.append({\u0026#39;Numbers\u0026#39;:df_china_lately.iloc[0][\u0026#39;Recovered\u0026#39;],\u0026#39;Status\u0026#39;:\u0026#39;Recovered\u0026#39;},ignore_index=True) df_china_status=df_china_status.append({\u0026#39;Numbers\u0026#39;:df_china_lately.iloc[0][\u0026#39;Deaths\u0026#39;],\u0026#39;Status\u0026#39;:\u0026#39;Deaths\u0026#39;},ignore_index=True) #Visualizing graphs fig = px.bar(df_china_status, y=\u0026#39;Numbers\u0026#39;, x=\u0026#39;Status\u0026#39;,title=\u0026#39;Latest Status in China\u0026#39; , color=\u0026#39;Numbers\u0026#39;) fig.show() The number of confirmed cases surged in January last week and mid of February but from March's first week there is no sudden rise in confirmed cases, it’s kind of stable after that The number of Deaths is increasing throughout the month of Feb and March, it keep on increasing when the number of confirmed cases are not increasing much this may be because previously confirmed cases were not able to recover The number of recoveries is pretty good, from February last week there is a sudden rise in recovered cases and that’s why active cases are very less. China did a pretty good job of fighting back #Calculating Death,Active Cases and Recovery percentage death_percentage=round((df_china_lately.iloc[0][\u0026#39;Deaths\u0026#39;]/df_china_lately.iloc[0][\u0026#39;Confirmed\u0026#39;])*100) active_cases_percentage=round((df_china_lately.iloc[0][\u0026#39;ActiveCases\u0026#39;]/df_china_lately.iloc[0][\u0026#39;Confirmed\u0026#39;])*100) recovery_percentage=round((df_china_lately.iloc[0][\u0026#39;Recovered\u0026#39;]/df_china_lately.iloc[0][\u0026#39;Confirmed\u0026#39;])*100) print(\u0026#34;In China Active Cases and Recovery percentage is {} ,{} respectively and Death rate is {} \u0026#34;.format(active_cases_percentage,recovery_percentage,death_percentage)) The rate of Active and Recovered Cases are 4% and 92% respectively and the Death rate is also 4% in China till 27th March 2020\nWell, this is a tremendous comeback for China. After seeing a rapid surge in confirmed cases Chinese government was able to control this pandemic in a short period. They are following an ideal graph for Active and Recovered cases, any country that is fighting COVID-19 should exhibit one of these characteristics,\nLet’s start looking at the trend in Italy\nNote: I will not mention code for other countries as the codes are identical\nThere is a sudden surge in confirmed cases in and after the first week of March The rate by which confirmed cases are increasing is way higher than the rate of recovery and deaths, which is not the case in China Within 10 days (between 12th-21st March) death tolls increased from 827 to 4825 The rate of Active and Recovered Cases are 77% and 13% respectively and the Death rate is also 11% in Italy till 27th March 2020\nLet\u0026rsquo;s check what happened in South Korea\nConfirmed cases have increased from 31 to 602 within 5 days i.e between February 19 and February 24, this may be because of Patient#31 of South Korea (refer to this article to read about this) From 15th March rate of confirmation is kind of stable and not increased much as compared to previous weeks After 15th March rate of recovery has increased drastically and the number of deaths is less as compared to confirmed cases The rate of Active and Recovered Cases are 50% and 49% respectively and the Death rate is also 1% in South Korea till 27th March 2020\nLike China, South Korea also did a commendable job in controlling this pandemic which is evident with their low death rate but the number of active cases is high which is close to 60%. But unlike Italy, despite of sudden surge in cases, they were able to handle them efficiently.\nLet’s list down the points which would be ideal in the current scenario\nThe rate of Confirmed Cases should not be increasing drastically The rate of Death should be very low The rate of Active cases should be gradually decreasing The rate of Recovery should be gradually increasing It’s evident from the above graphs that China and South Korea are exhibiting all these points. Any country who are in the early stages should exhibit these trends to fight against COVID-19\nTrends in USA and INDIA There is a huge surge in confirmed cases in the US after March 10, cases rises from 267 to 15,793 in 12 days(10 March-22 March) Death rates increased drastically from 1 to 117 within 22 days (1 March- 22 March) The rate of recovery was kind of flat till 15 March after that it increased The majority of cases are still active The rate of Active and Recovered Cases are 98% and 1% respectively and the Death rate is also 2% in the USA till 27th March 2020\nLet’s check about India\nIn India also the number of confirmed cases increased drastically from the first week of March, within 20 days it raised from 5 to 396 (2 March-22 March) but the numbers are significantly low as compared to US Death and Recovery rate is relatively low India also has huge number of Active cases The rate of Active and Recovered Cases are 90% and 8% respectively and the Death rate is also 2% in India till 27th March 2020\nThe USA and India are not doing well in terms of cumulative numbers and graphs, they have a high number of active cases and a low number of recoveries, which is not very encouraging\nChina, South Korea, and Italy Vs USA and India We will compare confirmed cases and Active Cases in these countries across weeks to find out where US and India can land in the upcoming weeks\nFor China, in initial weeks there is a surge in Confirmed Cases but after 4th week of reporting initial cases(#8 at the x-axis in the above graphs) active cases started reducing Similarly, for South Korea after the 4th week of reporting initial cases(#11 at the x-axis in the above graph) confirmed cases are not increasing much, and active cases are declining slowly For US and Italy after 7th week of reporting initial cases(#11 at the x-axis in the above graph) there is a sudden surge in confirmed cases and this is rapidly increasing week by week (We can see this by observing each Week Steps Height in the above graphs , which is quite high for US/Italy as compare to China/South Korea) As of now numbers are not high in India but here also there is a surge after 7th week of reporting initial cases(#11 at the x-axis in the above graph),this is significantly very less as compared to Italy and US Last Word Well, this is evident from the above analysis that China and South Korea can control and tackle COVID-19 efficiently as they have started all their measures early. Their active cases are coming down significantly after the 4th week of initially reported cases.\nWhereas Italy and the US are following the Polynomial Regression curve in Confirmed and Active cases which means in upcoming weeks it’s going to increase for sure. There is very little chance of a decline in this curve, especially for the US because there is no complete country lockdown and the US government is only focusing on areas where the number of cases is more.\nAlso, there was a significant amount of time between the initial cases and the sudden surge (close to 7 weeks) but there were no proper measures and actions during this period and I strongly believe this may be the reason behind the sudden surge because, in absence of lockdown, this time frame is sufficient enough to spread the virus.\nAs of now, fortunately, this is not happening for India as numbers are still less and if Indians follow 21 days of complete lockdown (starting from 24 March till 14th April) rigorously then the situation could be better, hope for the best.\nThoughts and Prayers with all who are fighting COVID-19 !!\nDisclaimer: Just to let you know, this blog post was originally published on Medium. If you\u0026rsquo;d like to check out the original, you can find it at this link.\n","permalink":"https://blogsbykush.com/covid19-insights-data-visualization/","summary":"In this post, the author explores the impact of COVID-19 on different countries using data visualization techniques. The focus is on countries like China, South Korea, Italy, the USA, and India. The analysis highlights trends in confirmed cases, active cases, recoveries, and deaths, shedding light on the effectiveness of measures taken by each country. The post emphasizes the importance of early action and proper containment strategies in controlling the spread of the virus","title":"COVID-19-Insights with Data Visualization"},{"content":"Introduction Regression testing is the common task of retesting software that has been changed or extended by new features during software development and most of the time retesting the whole program is not feasible with reasonable time and cost, and to overcome only a subset of all test cases is executed for regression testing, e.g., by executing test cases according to test case prioritization.\nThere are a vast amount of methods for test case selection exist but mostly it is a based on domain expertise of Test Engineers/Subject Matter Expert.As obvious, this manual process is time consuming, iterative and largely depends upon the engineer\u0026rsquo;s skills which mean there are high chances of missing some relevant test cases.\nIn this blog, I am explaining a proof of concept(POC) in which the selection of manual regression test cases is happening with the help of the Classification Learning model. In this approach I have considered metadata related to test cases and Natural Language test case descriptions as an input to classification learning models to predict the selection of test cases\nBelow image will summarize the proposed solution\nData Collection and Preparation For POC we have considered microservices Test cases across four release cycles as our Test Data.\nThis is an authorization microservice it is based on Oauth2 standard which is widely used in industry for authorization across systems/microservices. Refer this link for more information on Oauth2\nBefore going forward let\u0026rsquo;s understand how currently test selection process is happening in HP Cloud Print Platform\nThis image is specific to a particular organization but across the industry, the process is somewhat similar to this Release Manifest: It is collection of versioned stuff that is being deployed, configuration settings, and issues/stories/artifacts description which are going to be deployed in particular release.\nJIRA: Agile Project Management and Bug Tracking Tool.\nService Functionality Mapping File: Matrix which consists of mapping between Microservices and functionality , this will help users to understand impacted area when particular microservice getting affected.\nTest Rail: Test Management Tool.\nLet\u0026rsquo;s understand the process step by step\nStep 1: For every release Subject Matter Expert(SME) refer Release Manifest to understand which microservices are under test and details of the fixes/commits in that particular release\nStep 2: SME will take stories and Bug ID\u0026rsquo;s from Release Manifest and navigate to JIRA to get more relevant details , also based on the domain knowledge SME will refer Service Functionality Mapping File to understand impacted functionality\nStep 3: Based on the information from JIRA ,SME will again refer Service Functionality Mapping File to get the list of impacted fucntionality\nStep 4: Based on the impacted functionality list from above two steps SME will now navigate to Test Rail to search relevant test cases\nStep 5: With the help of domain knowledge and data collected from above steps SME will select the list of test cases from Test Rail\nLet’s take a quick look at test data which consists of regression test cases across four release cycles\nLet\u0026rsquo;s understand the columns/features present in dataset\nID: Unique Identifier of Records\nReleaseID : Release Identification number , Ex: R20.2.1 stands for \u0026lsquo;First release of 2 month of Year 2020\nType of Test Case : Cateogarization of Test cases , Ex: \u0026lsquo;Sanity\u0026rsquo; test cases are supposed to be executed for Sanity of microservice and \u0026lsquo;API/Functionality\u0026rsquo; test cases are for core functionality of microservice\nTestCaseTitle : Title or summary of test case\nTestCaseDescription : Steps for a test cases in Behavior Driven Devlopment(BDD) format\nError Prone Test Cases : Test cases which are covering high error prone area , these test cases must be executed in every release\nAutomation Status : Wheather test case is automated or not\nAny Defect : If there is any defect in the release manifest then it will be mapped to the corresponding/relevant test cases and this column will be marked as \u0026lsquo;Yes\u0026rsquo;\nJIRA Bug ID : Corresponding Bug ID of JIRA\nBug Description : JIRA Title/description of corresponding bug\nGIT Commit Message : For particular release , if there are any commits in GIT then the corresponding/relevant test cases and this column will be marked with commit messages\nTarget : Binary classification of Test Cases selection\nIt\u0026rsquo;s quite possible that in actual , SME/Test Engineer is not considering above features/columns for test cases selection. But we strongly believe that these features should be considered during test case selection as they are directly or indirectly impacting Release and Qualification Cycles\nExploratory Data Analysis Let\u0026rsquo;s explore these features one by one and it\u0026rsquo;s impact on target variable , in other words let\u0026rsquo;s try to understand how these features are related to selection of test cases. This analalysis will eventually help us in training our classifier model.\nLet\u0026rsquo;s start with Categorical Variables ,will begin with ReleaseID.\nWe will check that how test cases are selected across different releases\n#Getting unique value of releaseID column dataset[\u0026#39;ReleaseID\u0026#39;].unique() dataset[\u0026#39;ReleaseID\u0026#39;].value_counts() R20.1.1 166 R19.12.1 166 R20.2.1 166 R20.1.2 166 Name: ReleaseID, dtype: int64 Test data comprise of four different releases and for each release there are equal number of test cases , but not all of them are selected for execution.\nLet\u0026rsquo;s visualize that how many test cases selection is happening across release\nIt\u0026rsquo;s evident that selection of test cases is not uniform across the release and which is obvious , but we need to identify that what are the fetaures which play role in test case selection.\nLet\u0026rsquo;s see how defects are mapped across releases\nIn R20.1.1 release there are 18 bugs and in R20.2.1 release there are 3 bugs , from the previous graph we can say that maximum number of selected test cases are in these two releases only .\nWith this analysis we can say that Any Defect feature is playing role in selection of test case.\nNow let\u0026rsquo;s start exploring column Error Prone Test Cases\nNumber of Error Prone Test Cases is same in all release , on closer look it seems to be obvious because every release have similar set of test cases and these cases will be marked as \u0026lsquo;Error Prone\u0026rsquo; based on domain expertise of SME.\nLet\u0026rsquo;s see how these test cases are selected\nprintmd(\u0026#39;**Number of Not Selected Error Prone Test Cases are {}**\u0026#39;.format(len(dataset[(dataset[\u0026#39;Error Prone Test Cases\u0026#39;] == \u0026#39;Yes\u0026#39;) \u0026amp; (dataset[\u0026#39;Target\u0026#39;]==\u0026#39;No\u0026#39;)].index))) Number of Not Selected Error Prone Test Cases are 0\nThis means all Error Prone Test cases from each release are selected for executed , it means these cases are must executed test cases and this feature directly respobsible for selecting test cases\nLet\u0026rsquo;s put our focus on another features Automation Status and Type of Test Case\nFirst we will see what type of test cases are selected for executed\n#Calculating total type of test cases printmd(\u0026#34;**Test Cases Segregation \\n{}**\u0026#34;.format(dataset[\u0026#39;Type of Test Case\u0026#39;].value_counts())) Test Cases Segregation:\nAPI/Functionality Integration Sanity 640 12 12 All Sanity and Integration type of test cases are selected in all four release. With this we can conclude that Type of Test Cases are directly related to selection of test cases.\nNow we will check Automation Status feature , will try to understand that how test cases are selected across releases based on automation status\nTotal Number of Automated Test Cases are 236 Total number of selected test cases which are automated 139 In our data set we have total 236 automated test cases and out of which 139 are selected test cases. Number of automated test cases are same in all four releases which is 59 .\nAbove data tell us that though number of automation test cases are same across all releases , selection of the same may depends upo other factors.\nLet\u0026rsquo;s find out more how many automated test cases are selected in different releases\nHighest number of automated selected test cases i.e 42% , 41% are in R20.1.1 and R20.2.1 releases respectively and from our pevious analysis we can say that these are the two releases where we have maximum number of bugs and highest number selected test cases.\nWith this it\u0026rsquo;s evident that Automation Test Cases feature played a role in test case selection specially when we have lot to test.\nAbove EDA looks obvious to many but the intent is to establish a relationship between features and target variable, so that classifier can learn the same pattern\nWith this we are done with analysis of all categorical data columns present in our data set and all of them looks related to our Target Variable , we will keep all of them in our further analysis.\nLet\u0026rsquo;s start with Textual columns , we have following Text columns\nTest Case Title Test Case Description JIRA Bug ID Bug Description GIT Commit Message From above features list we can remove \u0026lsquo;JIRA Bug ID\u0026rsquo; from our analysis because this is just a alphanumeric respresentation of Bugs and more details are present in next feature \u0026lsquo;Bug Description\u0026rsquo;.\nAlso , in our data set \u0026lsquo;GIT commit Message\u0026rsquo; is very inconsistent and most of the commit messages are of bug fixing and for these fixes commit messages are either bug ID or bug description which is kind of duplicate of data present in feature \u0026lsquo;Bug Description\u0026rsquo;. We can drop GIT commit message column also.\nIt’s recommended that ‘Commit Messages’ should be there for test cases selection but for this relevant messages should enter at the time of commit and same has to be captured properly in Release Manifest for each releases\nLet\u0026rsquo;s start looking at \u0026lsquo;Test Case Title\u0026rsquo; and \u0026lsquo;Test Case Description\u0026rsquo;\nThere are no empty records in Test Case Title but same is not true for Description , majority of the description is empty.\nMoreover , if we take a closer look in Title then we can say that it consist of summary of a test case which itself is a good information to select or unselect the test case. With this information we can drop \u0026lsquo;Test Case Description\u0026rsquo; from our list.\nOur Textual feature list for analysis will be reduced to following\nTest Case Title Bug Description \u0026lsquo;Test Case Title\u0026rsquo; should be primary entry point for test cases selection as it give summary/intention behind the test case and this makes perfect choice for Test Case Selection\nWe have already concluded that \u0026lsquo;Any Defect\u0026rsquo; column has direct relation with test case selection and if \u0026lsquo;Any Defect\u0026rsquo; column has a value then corresponding \u0026lsquo;Bug Description\u0026rsquo; should be there. With this co relation we can say that this column should be there in our feature list but in our data set only few entries are there fo \u0026lsquo;Bug Description\u0026rsquo; , so we can skip this column also from our list\nprintmd(\u0026#34;**Number of records with bug descriptions are {}**\u0026#34;.format(len(dataset[pd.notnull(dataset[\u0026#34;Bug Description\u0026#34;])].index))) Number of records with bug descriptions are 21\n#Drop not required column dataset=dataset.drop([\u0026#39;TestCaseDescription\u0026#39;,\u0026#39;GIT Commit Message\u0026#39;,\u0026#39;JIRA Bug ID\u0026#39;,\u0026#39;Bug Description\u0026#39;],axis=1) dataset.head() Id ReleaseID Type of Test Case TestCaseTitle Error Prone Test Cases Automation Status Any Defect Target 0 1 R20.2.1 Sanity Get the short and detailed health status APIs No Yes Yes 1 1 2 R20.2.1 Sanity Get the AuthZ service metadata No Yes Yes 1 2 3 R20.2.1 Sanity Get the public keys for validating token No Yes Yes 1 3 4 R20.2.1 API/Functionality Verify Client delegation API: Exchange Access ... Yes Yes No 1 4 5 R20.2.1 API/Functionality Verify Client delegation API:Exchange Access t... No Yes No 1 Now, we have one text column and five categorical columns and based on our analysis we can say that all of these are related to selection of test cases i.e. related to \u0026lsquo;Target\u0026rsquo; variable.\nWe can freeze above list of features for training classifier models , But before that we need to convert all categorical variables to encoded variables form and text columns into sparse matrix for features\n#Encoding all categorical variable from sklearn.preprocessing import LabelEncoder, OneHotEncoder labelencoder = LabelEncoder() dataset[\u0026#39;ReleaseID\u0026#39;]=labelencoder.fit_transform(dataset[\u0026#39;ReleaseID\u0026#39;]) dataset[\u0026#39;Error Prone Test Cases\u0026#39;]=labelencoder.fit_transform(dataset[\u0026#39;Error Prone Test Cases\u0026#39;]) dataset[\u0026#39;Type of Test Case\u0026#39;]=labelencoder.fit_transform(dataset[\u0026#39;Type of Test Case\u0026#39;]) dataset[\u0026#39;Automation Status\u0026#39;]=labelencoder.fit_transform(dataset[\u0026#39;Automation Status\u0026#39;]) dataset[\u0026#39;Any Defect\u0026#39;]=labelencoder.fit_transform(dataset[\u0026#39;Any Defect\u0026#39;]) dataset.head() Id ReleaseID Type of Test Case TestCaseTitle Error Prone Test Cases Automation Status Any Defect Target 0 1 3 2 Get the short and detailed health status APIs 0 1 1 1 1 2 3 2 Get the AuthZ service metadata 0 1 1 1 2 3 3 2 Get the public keys for validating token 0 1 1 1 3 4 3 0 Verify Client delegation API: Exchange Access ... 1 1 0 1 4 5 3 0 Verify Client delegation API:Exchange Access t... 0 1 0 1 With this all categorical columns are encoded into numerical values. Now , let\u0026rsquo;s deal with our only Text column i.e. \u0026lsquo;TestCaseTitle\u0026rsquo;.\nFor this we will apply Natural Language Processing (NLP) to convert text data into features for classifier model\nLet\u0026rsquo;s create Corpus from \u0026lsquo;TestCaseTitle\u0026rsquo; , Corpus is a simplified version of our test case title data that contain clean text data.\nTo create Corpus we have to perform the following actions\nRemove unwanted words: Removal of unwanted words such as special characters and numbers to get only pure text. We will do it by specify our pattern using re library Transform words to lowercase: Transform words to lowercase because upper and lower case have diffirent ASCII codes Remove stopwords:Stop words are usually the most common words in a language and they will be irrelevant in determining the nature Stemming words:Stemming is the process of reducing words to their word stem, base or root form. We use stemming to reduce Bag of Words dimensionality #Importing required libraries import re import nltk from nltk.corpus import stopwords from nltk.stem.porter import PorterStemmer corpus_title = [] pstem = PorterStemmer() for i in range(dataset[\u0026#39;TestCaseTitle\u0026#39;].shape[0]): #Remove unwanted words text = re.sub(\u0026#34;[^a-zA-Z]\u0026#34;, \u0026#39; \u0026#39;, dataset[\u0026#39;TestCaseTitle\u0026#39;][i]) #Transform words to lowercase text = text.lower() text = text.split() #Remove stopwords then Stemming it text = [pstem.stem(word) for word in text if not word in set(stopwords.words(\u0026#39;english\u0026#39;))] text = \u0026#39; \u0026#39;.join(text) #Append cleaned tweet to corpus corpus_title.append(text) printmd(\u0026#34;**Corpus created successfully**\u0026#34;) Corpus created successfully\nFrom above Corpus we will create Bag of Words , which is representation of text that describes the occurrence of words within a document. It involves two things:\nA vocabulary of known words A measure of the presence of known words This is called Bag of Word because any information about the order or structure of words in the document is discarded and the model is only concerned with whether the known words occur in the document, not where they occur in the document.\nFor Example :\nWe can do this using scikit-learn\u0026rsquo;s CountVectorizer, where every row will represent a different test case and every column will represent a different word.\nCountvectorizer converts a collection of text documents to a matrix of token counts. It is important to note here that CountVectorizer comes with a lot of options to automatically do preprocessing, tokenization, and stop word removal.However, i did all the process manually above to just get a better understanding.\n# Creating the Bag of Words model from sklearn.feature_extraction.text import CountVectorizer cv = CountVectorizer() text_vectors= cv.fit_transform(corpus_title).toarray() #Convert text vectors into data frame text_vectors_df=pd.DataFrame(text_vectors) printmd(\u0026#34;**Dimension for Text features are {}**\u0026#34;.format(text_vectors_df.shape)) Dimension for Text features are (664, 115)\n#Getting Target variable into Y variable y=dataset[[\u0026#39;Target\u0026#39;]].values #Converting 2 dimensional y and y_pred array into single dimension y=y.ravel() #Removing \u0026#39;Target\u0026#39; and \u0026#39;TestCaseTitle\u0026#39; columns from actual dataset dataset=dataset.drop([\u0026#39;Target\u0026#39;,\u0026#39;TestCaseTitle\u0026#39;],axis=1) #Creating new data frame with all categorical feature and Text features for training classifier models X=pd.concat([dataset,text_vectors_df],axis=1).values printmd(\u0026#34;**Dimension for features data frame are {}**\u0026#34;.format(X.shape)) #X.head() Dimension for features data frame are (664, 121)\nWith this we have completed our EDA and Feature Engineering , let\u0026rsquo;s start with classifier models creation\nLearning and Classification Now we will build our models, for current data set we are using follwoing models\nLogistic Regression Model Gaussian Naive Bayes Model Multinomial Naive Bayes Model If we have large data set we can use following models but for current data set we are avoiding them\nDecision Tree Model Gradient Boosting Model K - Nearest Neighbors Model After building models we will evaluate our model based on confusion matrix , which is formed from the four outcomes produced as a result of binary classification\nA binary classifier predicts all data instances of a test dataset as either positive or negative. This classification (or prediction) produces four outcomes — true positive, true negative, false positive and false negative.\nTrue positive (TP): correct positive prediction False positive (FP): incorrect positive prediction True negative (TN): correct negative prediction False negative (FN): incorrect negative prediction A confusion matrix of binary classification is a two by two table formed by counting of the number of the four outcomes of a binary classifier. We usually denote them as TP, FP, TN, and FN instead of “the number of true positives”, and so on\nError rate (ERR) is calculated as the number of all incorrect predictions divided by the total number of the dataset. The best error rate is 0.0, whereas the worst is 1.0\nAccuracy (ACC) is calculated as the number of all correct predictions divided by the total number of the dataset. The best accuracy is 1.0, whereas the worst is 0.0. It can also be calculated by 1 – ERR\nF1 Score\nBased on above formulas we can evaluate our models in Train and Test data set\nRegression Classifier F1 Score is 68.6 MultinomialNB Classifier F1 Score is 70.1 GaussianNB Classifier F1 Score is 70.1 Last Word Prediction could be improved if we have large data set , becuase then we can apply models like Decision Tree Model ,Gradient Boosting Model and K-Nearest Neighbors Model . We have checked these models for larger data set and prediction was pretty good Training data can be increase by adding more releases data Above approach can be used as a reference for similar problem statement Note : I have captured very less code snippet for this blog , if interested in actual work then refer this notebook\nDisclaimer: Just to let you know, this blog post was originally published on Medium. If you\u0026rsquo;d like to check out the original, you can find it at this link.\n","permalink":"https://blogsbykush.com/regression-test-case-selection-using-ml/","summary":"Regression testing plays a crucial role in software development, but retesting the entire program after making changes or adding new features is often impractical. To address this, a subset of test cases is executed for regression testing. In this blog post, we explore a proof of concept (POC) that uses a Classification Learning model to assist in the selection of manual regression test cases. By considering metadata and natural language descriptions of test cases, the model predicts which test cases should be selected. The post also covers data collection and preparation, as well as exploratory data analysis to understand the relationship between various features and test case selection","title":"Regression Test Case Selection Using ML"},{"content":"Introduction Exploratory data analysis (EDA) is an approach to analyzing data sets to summarize their main characteristics, often with visual methods. A statistical model can be used or not, but primarily EDA is for seeing what the data can tell us beyond the formal modeling or hypothesis testing task. When I started my journey in the Data Science field I always had difficulty with the starting point of any problem but after reading a few exemplary Kernels in Kaggle I have realized the power of Exploratory Data Analysis and its impact on Data Modeling and Predictions\nI am trying to explain how we can do EDA and Feature Engineering as the simplest way to get some insight into the Titanic Disaster. I have put only specific code snippets before each visualization and analysis, if anyone is interested in full code then refer to the link provided at the end.\nThe sinking of the Titanic is one of the most infamous shipwrecks in history. On April 15, 1912, during her maiden voyage, the widely considered “unsinkable” RMS Titanic sank after colliding with an iceberg. Unfortunately, there weren’t enough lifeboats for everyone on board, resulting in the death of 1502 out of 2224 passengers and crew.\nI have used dataset which is provided by Kaggle for Titanic: Machine Learning from Disaster Competition\nFeatures Analysis Let’s import required libraries for EDA\n#Importing required libraries import numpy as np import seaborn as sns import pandas as pd import matplotlib.pyplot as plt import plotly.graph_objects as go import plotly.express as px from sklearn.ensemble import RandomForestClassifier #Importing train data set ds_train=pd.read_csv(\u0026#34;/\u0026lt;InputDirectory\u0026gt;/train.csv\u0026#34;) #Checking features in train data set ds_train.head() Now , analyze these features/variables one by one\nSurvived is a target variable where the survival of a passenger is predicted in binary format i.e. 0 for Not Survived and 1 for Survived\nPassengerId and Ticket variables can be assumed as Random unique Identifiers of Passengers and they don\u0026rsquo;t have any impact on survival, hence we can ignore them\nPclass is an ordinal datatype for the ticket class, it can be considered as the passenger\u0026rsquo;s Socio-Economic Status and it may impact the passenger’s survival chances so we will keep this in our analysis. It\u0026rsquo;s unique values are 1 = Upper Class, 2 = Middle Class and 3= Lower Class\nName is self-explanatory, we will skip this variable from our analysis\nSex or Gender could have played an important role in survival because during any evacuation from disaster, preference will be given to the female gender and to test this notion we will consider gender in our analysis\nSibSp and Parch represent the total number of the passenger\u0026rsquo;s siblings/spouse and parents/children on board respectively, they could be used to create a new variable called \u0026lsquo;Family Size\u0026rsquo; (Creating a new feature/variable is an example of Feature Engineering)\nAge could have also played a role in survival, so we will keep this in our feature list\nFare is also an indicator of the Socio-Economic Status of passengers, let\u0026rsquo;s keep this in our feature list\nCabin is the Cabin number of the passenger and it can be used in Feature engineering to get an approximate position of the passenger when the accident happened, also from deck level, we can deduce Socioeconomic status. However, after looking at the data it looks like there are many null values so we can drop this column from our feature list.\nEmbarked is a port of embarkation of passengers and this may have an impact on the target variable so we will keep this variable for now. It has 3 unique values , C = Cherbourg ,Q = Queenstown and S = Southampton\nVisualization Now, we will try to see the relation between selected features by creating Seaborn and Plotly visualization\nFirst, start with the passenger’s Age\n#Converting Age into series and visualizing the age distribution age_series=pd.Series(ds_train[\u0026#39;Age\u0026#39;].value_counts()) fig=px.scatter(age_series,y=age_series.values,x=age_series.index) fig.update_layout( title=\u0026#34;Age Distribution\u0026#34;, xaxis_title=\u0026#34;Age in Years\u0026#34;, yaxis_title=\u0026#34;Count of People\u0026#34;, font=dict( family=\u0026#34;Courier New, monospace\u0026#34;, size=18, ) ) fig.show() We can deduce a few points from the above graph\nMajority of passengers aged more than 20 years and less than 50 years 30 passengers share the same age i.e. 24 years 164 passengers share the same age Let’s check how Gender is distributed among passengers\nprint(\u0026#34;Number of Passengers Gender Wise \\n{}\u0026#34;.format(ds_train[\u0026#39;Sex\u0026#39;].value_counts())) #Gender wise distribution fig = go.Figure(data=[go.Pie(labels=ds_train[\u0026#39;Sex\u0026#39;],hole=.4)]) fig.update_layout( title=\u0026#34;Sex Distribution\u0026#34;, font=dict( family=\u0026#34;Courier New, monospace\u0026#34;, size=18 )) fig.show() It’s quite evident that the number of male passengers is almost double of female passengers.\nLet’s see how many females and males survived across different age groups.\n#Create categorical variable graph for Age,Sex and Survived variables sns.catplot(x=\u0026#34;Survived\u0026#34;, y=\u0026#34;Age\u0026#34;, hue=\u0026#34;Sex\u0026#34;, kind=\u0026#34;swarm\u0026#34;, data=ds_train,height=10,aspect=1.5) plt.title(\u0026#39;Passengers Survival Distribution: Age and Sex\u0026#39;,size=25) plt.show() Well it’s pretty evident from the above graph that the majority of female passengers are survived\nMajority of Male passengers aged between 20 to 50 years had not survived. It means most of the young men had not survived this disaster Oldest male passenger aged 80 years,had survived Age and Sex were major factors in deciding the passenger’s fate Now, let’s see Pclass variable relation with survival\n#Visualize relation between Pclass and Survival fig = go.Figure(data=[go.Pie(labels=ds_train[\u0026#39;Pclass\u0026#39;],hole=.4)]) fig.update_layout( title=\u0026#34;PClass Distribution\u0026#34;, font=dict( family=\u0026#34;Courier New, monospace\u0026#34;, size=18 )) fig.show() More than half of the passengers were traveling in Lower Class.\nLet’s see how survival is linked with Pclass\n#Visualize PClass and Survival #Create categorical variable graph for Age,Pclass and Survived variables sns.catplot(x=\u0026#34;Survived\u0026#34;, y=\u0026#34;Age\u0026#34;, hue=\u0026#34;Pclass\u0026#34;, kind=\u0026#34;swarm\u0026#34;, data=ds_train,height=10,aspect=1.5) plt.title(\u0026#39;Passengers Survival Distribution: Age and Pclass\u0026#39;,size=25) plt.show() Again , majority of young male passengers aged between 20 to 50 years and travelling in lower class had not survived Oldest male passenger who survived the disaster was travelling in upper class Young men who survived the disaster were travelling in upper class If the passenger was a man aged between 20–50 years, and not so rich at the time of travel then their chances of survival were very less\nTo support our Socio-Economic Status theory let’s focus on one more variable Fare\n#Visualize Fare and Survival #Create categorical variable graph for Sex,Fare and Survived variables sns.catplot(x=\u0026#34;Survived\u0026#34;, y=\u0026#34;Fare\u0026#34;, hue=\u0026#34;Sex\u0026#34;, kind=\u0026#34;swarm\u0026#34;, data=ds_train,height=10,aspect=1.5) plt.title(\u0026#39;Passengers Survival Distribution: Fare and Sex\u0026#39;,size=25) plt.show() In the above graph, for the feature ‘Sex’ consider 1 for females and 0 for males. It’s evident that female passengers with lower ticket fares survived the disaster and a few male passengers with the highest fare also survived.\nIt means when it comes to gender, the female got preference across all the classes otherwise Socio-Economic Status played an important role in survival.\nNow, we will see Embarked variable’s impact on survival\n#Visualize relation between Embarked and Survival fig = go.Figure(data=[go.Pie(labels=ds_train[\u0026#39;Embarked\u0026#39;],hole=.4)]) fig.update_layout( title=\u0026#34;Embarked Distribution\u0026#34;, font=dict( family=\u0026#34;Courier New, monospace\u0026#34;, size=18 )) fig.show() The majority of passengers embarked from Southampton, let’s visualize its survival distribution.\n#Visualize Embarked and Survival #Create categorical variable graph for Embarked,Age and Survived variables sns.catplot(x=\u0026#34;Survived\u0026#34;, y=\u0026#34;Age\u0026#34;, hue=\u0026#34;Embarked\u0026#34;, kind=\u0026#34;swarm\u0026#34;, data=ds_train,height=10,aspect=1.5) plt.title(\u0026#39;Passengers Survival Distribution: Embarked and Age\u0026#39;,size=25) plt.show() We can not deduce any direct relation between Embarked and Survival.\nLet’s check the correlation coefficient between these features\n# Training set high correlations ds_train.corr() We can see a direct correlation between the ‘Survived’ and ‘Fare’ variables, other variables are in-directly related to Survival\nAge is correlated to Fare and Fare is correlated to Survived and our analysis also shows how Age played a role in survival, by this we can say that Age is related to Survival SibSp and Parch are related to each other and also both are related to Fare which makes sense because more people means more fare, by virtue of this both can be related to Survived Feature Engineering Feature engineering is the process of using domain knowledge to extract features from raw data via data mining techniques. These features can be used to improve the performance of machine learning algorithms. Having and engineering good features will allow you to most accurately represent the underlying structure of the data and therefore create the best model.\nFeatures can be engineered by decomposing or splitting features, from external data sources, or aggregating or combining features to create new features.\nLet’s start Feature Engineering by creating a new variable Family Size by adding SibSp, Parch, and One(Current Passenger)\n#Add new column \u0026#39;Family Size\u0026#39; in training model set ds_train[\u0026#39;Family_Size\u0026#39;] = ds_train[\u0026#39;SibSp\u0026#39;] + ds_train[\u0026#39;Parch\u0026#39;] + 1 print(\u0026#34;Family Size column created sucessfully\u0026#34;) ds_train.head() Now we will see how the Family size will is related to Survived variable\n#Visualize Family size and Survival sns.barplot(x=\u0026#34;Family_Size\u0026#34;, y=\u0026#34;Age\u0026#34;, hue=\u0026#34;Survived\u0026#34;, data=ds_train,palette = \u0026#39;rainbow\u0026#39;) plt.title(\u0026#39;Family Size - Age Survival Distribution\u0026#39;,size=20) plt.show() sns.catplot(y=\u0026#34;Family_Size\u0026#34;, x=\u0026#34;Survived\u0026#34;, hue=\u0026#39;Sex\u0026#39;,kind=\u0026#34;swarm\u0026#34;, data=ds_train,height=8,aspect=1.5) plt.title(\u0026#39;Family Size - Gender Survival Distribution\u0026#39;,size=25) plt.show() Chances of survival are less for large Families (\u003e5 members) If the family size is small then the main passenger’s gender decides on survival, this supports the previous deduction of Gender’s role in the survival Note: Survival data is marked for main passengers and not for the whole family, whereas family members’ names must be there in the list and they may or may not be survived. In other words, by just looking at the survival column we can not deduce that the fate of all family members was the same\nLast Word We can see that by just visualizing the relation between a few variables we got so many insights and further we can use this newly gained knowledge regarding a feature in training data models by adding new features and removing the unnecessary ones.\nRefer to Kaggle Kernel or Juypter Notebook for whole analysis and data modeling\nDisclaimer: Just to let you know, this blog post was originally published on Medium. If you\u0026rsquo;d like to check out the original, you can find it at this link.\n","permalink":"https://blogsbykush.com/guidefornewbie-eda-featureeng-modeling-evaluation/","summary":"In this blog post, the author provides a beginner\u0026rsquo;s guide to exploratory data analysis (EDA) and feature engineering using the Titanic disaster dataset. They explain the importance of EDA in understanding data and its impact on data modeling and predictions. The post includes code snippets and visualizations to analyze various features such as age, gender, passenger class, fare, and embarked port. The author also discusses correlations between different variables and demonstrates the process of feature engineering by creating a new variable called Family Size. The post concludes by emphasizing the significance of EDA in gaining insights and improving data modeling","title":"Beginner's Guide to Exploratory Data Analysis and Feature Engineering"},{"content":" Hi, I\u0026rsquo;m Kush. I\u0026rsquo;m a Technical Product Manager in Bangalore, with 17+ years in tech and 5+ years in product management. I lead an enterprise AI platform by day and build my own AI products after hours.\nThis blog is where I learn in public, show what I build and explain it in plain words, for product people and non-technical folks in tech.\nWhat I work on I own the product vision for an internal AI platform that gives 10+ teams access to public and proprietary LLMs through one API, with the access controls, usage tracking and governance that make AI safe to use at enterprise scale. I also build AI agents for product work, covering the lifecycle from market research and business case to PRD, specification and Jira epics, with human review at every stage, which took each stage from weeks to about 2–3 days. Through workshops and office hours I\u0026rsquo;ve helped 100+ colleagues get started, and 5+ PMs in other teams now build and use their own agents.\nWhat I\u0026rsquo;m building After hours, I build AI products end to end: I find the problem, write the requirements, build with AI coding tools, ship, and see what actually works. I share the decisions, the mistakes and the lessons in the Build Log.\nHow this blog works I believe learning is a continuous loop: to truly learn something, you have to experiment with it and transform it into something of your own. This site is that loop, in public, built on three things I value: learning, writing and simplicity.\nLearn: Tech Digest, Digital Dhaba, my weekly AI digest every Thursday, and Learning Notes, my own notes on what I read. Build: Build Log, products I\u0026rsquo;m building in public. Explain: Concept Breakdown, comic-style breakdowns of AI and product concepts. No jargon. The learning loop, from Michael Simmons. Get in touch LinkedIn is the best place to reach me, or find me here:\nGet Digital Dhaba and my weekly letter by email: subscribe.\n","permalink":"https://blogsbykush.com/about/","summary":"\u003cimg class=\"about-photo\" src=\"/assets/images/kush-avatar.jpg\" alt=\"Kush Bhatnagar\" width=\"180\" height=\"180\"\u003e\n\u003cp\u003eHi, I\u0026rsquo;m Kush. I\u0026rsquo;m a \u003cstrong\u003eTechnical Product Manager\u003c/strong\u003e in Bangalore, with 17+ years in tech and 5+ years in\nproduct management. I lead an enterprise AI platform by day and build my own AI products after hours.\u003c/p\u003e\n\u003cp\u003eThis blog is where I learn in public, show what I build and explain it in plain words, for product people\nand non-technical folks in tech.\u003c/p\u003e\n\u003ch2 id=\"what-i-work-on\"\u003eWhat I work on\u003c/h2\u003e\n\u003cp\u003eI own the product vision for an internal AI platform that gives 10+ teams access to public and proprietary\nLLMs through one API, with the access controls, usage tracking and governance that make AI safe to use at\nenterprise scale. I also build AI agents for product work, covering the lifecycle from market research and\nbusiness case to PRD, specification and Jira epics, with human review at every stage, which took each stage\nfrom weeks to about 2–3 days. Through workshops and office hours I\u0026rsquo;ve helped 100+ colleagues get started,\nand 5+ PMs in other teams now build and use their own agents.\u003c/p\u003e","title":"About Kush"},{"content":"I learn best by building. This is where I document the products I build in public, from the first idea to launch: what I decided and why, what broke, and what I learned about building products along the way.\n","permalink":"https://blogsbykush.com/build-log/","summary":"\u003cp\u003eI learn best by building. This is where I document the products I build in public, from the first idea to launch: what I decided and why, what broke, and what I learned about building products along the way.\u003c/p\u003e","title":"Build Log"},{"content":"Every breakdown takes one concept from Machine Learning, GenAI, Cloud or Product Management (and sometimes life) and explains it through a short comic conversation, followed by a plain-words explanation and why it matters if you build products.\n","permalink":"https://blogsbykush.com/concept-breakdown/","summary":"\u003cp\u003eEvery breakdown takes one concept from \u003cstrong\u003eMachine Learning, GenAI, Cloud or Product Management\u003c/strong\u003e (and sometimes life) and explains it through a short comic conversation, followed by a plain-words explanation and why it matters if you build products.\u003c/p\u003e","title":"Concept Breakdown"},{"content":"When something I read changes how I think, I write a short note: what the source says, what I took from it, and what I\u0026rsquo;m still unsure about. Claude helps with a first research pass; the note and the takeaways are mine, checked against the original. Many of the sources come from Digital Dhaba.\n","permalink":"https://blogsbykush.com/learning-notes/","summary":"\u003cp\u003eWhen something I read changes how I think, I write a short note: what the source says, what I took from it, and what I\u0026rsquo;m still unsure about. Claude helps with a first research pass; the note and the takeaways are mine, checked against the original. Many of the sources come from \u003ca href=\"/tech-digest/\"\u003eDigital Dhaba\u003c/a\u003e.\u003c/p\u003e","title":"Learning Notes"},{"content":"In this series, I\u0026rsquo;ll be sharing my personal reflections on the books that have resonated with me, books that have left a lasting impression, sparked new insights or simply offered a moment of respite from the daily grind. My reading list is eclectic and random, but each book has something to offer\n","permalink":"https://blogsbykush.com/my-bookshelf-chronicles/","summary":"\u003cp\u003eIn this series, I\u0026rsquo;ll be sharing my personal reflections on the books that have resonated with me, books that have left a lasting impression, sparked new insights or simply offered a moment of respite from the daily grind. My reading list is eclectic and random, but each book has something to offer\u003c/p\u003e","title":"My Bookshelf Chronicles!!"},{"content":"Every Thursday you get Digital Dhaba, a weekly roundup of what\u0026rsquo;s happening in AI and tech and why it matters to people who build products. On Sundays, a short letter rounds up what I published that week: Concept Breakdowns, Build Log entries and Learning Notes.\nJoin the listTwo emails a week at most: Digital Dhaba on Thursday, and a Sunday letter with my new posts (only if there are any). No spam, unsubscribe anytime.\nYou'll get a confirmation email. Click it, and you're in.\n","permalink":"https://blogsbykush.com/subscribe/","summary":"\u003cp\u003eEvery Thursday you get \u003cstrong\u003eDigital Dhaba\u003c/strong\u003e, a weekly roundup of what\u0026rsquo;s happening in AI and tech and why it matters to people who build products. On Sundays, a short letter rounds up what I published that week: \u003cstrong\u003eConcept Breakdowns\u003c/strong\u003e, \u003cstrong\u003eBuild Log\u003c/strong\u003e entries and \u003cstrong\u003eLearning Notes\u003c/strong\u003e.\u003c/p\u003e\n\n\u003cdiv class=\"newsletter-signup\"\u003e\u003ch3\u003eJoin the list\u003c/h3\u003e\u003cp\u003eTwo emails a week at most: Digital Dhaba on Thursday, and a Sunday letter with my new posts (only if there are any). No spam, unsubscribe anytime.\u003c/p\u003e","title":"Subscribe"},{"content":"In this series, I dive into the landscape of MLOps tools that enable seamless integration, testing, and deployment of ML models. In addition, I explore the numerous benefits offered by various AWS cloud services specifically designed for Machine Learning lifecycles. Throughout this series, I will also delve into the complexities of designing effective model deployment architectures, while venturing into the realms of continuous delivery and automated machine learning pipelines\n","permalink":"https://blogsbykush.com/mlops-playground/","summary":"\u003cp\u003eIn this series, I dive into the landscape of MLOps tools that enable seamless integration, testing, and deployment of ML models. In addition, I explore the numerous benefits offered by various AWS cloud services specifically designed for Machine Learning lifecycles. Throughout this series, I will also delve into the complexities of designing effective model deployment architectures, while venturing into the realms of continuous delivery and automated machine learning pipelines\u003c/p\u003e","title":"The MLOps Playground"}]