<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Untitled Publication]]></title><description><![CDATA[Untitled Publication]]></description><link>https://ademuyiwaotubusin.hashnode.dev</link><generator>RSS for Node</generator><lastBuildDate>Sat, 10 Oct 2026 08:09:21 GMT</lastBuildDate><atom:link href="https://ademuyiwaotubusin.hashnode.dev/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[How I Built a Multimodal AI Maths Tutor with Gemini 2.5 Flash and Google Cloud Run]]></title><description><![CDATA[#GeminiLiveAgentChallenge | Built for the Gemini Live Agent Challenge Hackathon


Disclosure: I created this blog post for the purposes of entering the Gemini Live Agent Challenge hackathon. All code,]]></description><link>https://ademuyiwaotubusin.hashnode.dev/how-i-built-a-multimodal-ai-maths-tutor-with-gemini-2-5-flash-and-google-cloud-run</link><guid isPermaLink="true">https://ademuyiwaotubusin.hashnode.dev/how-i-built-a-multimodal-ai-maths-tutor-with-gemini-2-5-flash-and-google-cloud-run</guid><category><![CDATA[GeminiLiveAgentChallenge.]]></category><category><![CDATA[ai agents]]></category><category><![CDATA[hackathon]]></category><dc:creator><![CDATA[Ademuyiwa Otubusin]]></dc:creator><pubDate>Mon, 16 Mar 2026 21:31:37 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/64e360227fd57087b1184938/c311de61-c682-4558-ab54-0492e64bc95e.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>#GeminiLiveAgentChallenge</strong> | Built for the Gemini Live Agent Challenge Hackathon</p>
<hr />
<blockquote>
<p><strong>Disclosure:</strong> I created this blog post for the purposes of entering the <strong>Gemini Live Agent Challenge</strong> hackathon. All code, architecture decisions, and findings described here are from my own project, built during the hackathon period. When sharing on social media, please use the hashtag <strong>#GeminiLiveAgentChallenge</strong>.</p>
</blockquote>
<hr />
<p>Every student has been there. It's late, there's a maths exam tomorrow, and the only AI tool available just hands you the answer. You copy it, submit the homework, and learn absolutely nothing. The test comes and you're exposed.</p>
<p>That frustration was the starting point for <strong>Math Tutor AI Agent</strong> — a multimodal AI tutoring system built entirely on Google AI and Google Cloud during the Gemini Live Agent Challenge hackathon.</p>
<blockquote>
<p><em>"The tutor that teaches. Not the AI that answers."</em></p>
</blockquote>
<p>The live app runs at <code>ai-agent-frontend-1087118236338.us-central1.run.app</code>. The repositories are open source at <code>github.com/tay4real/ai-agent-backend</code> and <code>github.com/tay4real/ai-agent-frontend</code>.</p>
<hr />
<h2>The problem this solves</h2>
<p>Access to quality mathematics tutoring is one of the most unequally distributed educational resources in the world. Private tutors are expensive. Classroom teachers are stretched. And every AI tool available today shares the same fundamental flaw: they give students the answer instead of building their understanding.</p>
<p>The students who succeed in maths almost universally had access to someone patient enough to work through problems with them — explaining the reasoning, not just the result. A student in Lagos or Dhaka has the same intellectual potential as a student in London. What they often lack is that access.</p>
<p>Math Tutor AI Agent is built to close that gap. It is not a calculator. It is a tutor.</p>
<hr />
<h2>Why Gemini 2.5 Flash was the right model</h2>
<p>The first architectural decision was the AI model. I chose <strong>Google Gemini 2.5 Flash</strong> for three specific reasons.</p>
<h3>1. Native multimodality in a single API call</h3>
<p>Math Tutor AI Agent accepts problems via text, image, and voice. With Gemini 2.5 Flash, all three modalities flow through a single API call — no separate OCR service, no separate audio transcription model.</p>
<pre><code class="language-javascript">const response = await geminiModel.generateContentStream({
  contents: [{
    role: "user",
    parts: [
      { text: sessionContext },
      {
        inlineData: {
          mimeType: "image/jpeg",
          data: base64ImageData   // handwritten homework photo
        }
      },
      { text: "Guide me through this step by step." }
    ]
  }]
});
</code></pre>
<h3>2. Function calling for verified computation</h3>
<p>Gemini's function calling API lets the agent autonomously invoke <code>search()</code> for concept lookups and <code>calculator()</code> for symbolic computation — without any explicit routing logic in application code.</p>
<h3>3. Cost and speed profile</h3>
<p>For an always-available tutoring system, response latency and cost per token matter. Gemini 2.5 Flash is the fastest model in the Gemini family and the most cost-efficient for this use case.</p>
<hr />
<h2>System architecture</h2>
<p>The system is two repositories: a Node.js backend and a React frontend, communicating via REST and WebSocket.</p>
<table>
<thead>
<tr>
<th>Layer</th>
<th>Technology</th>
<th>Role</th>
</tr>
</thead>
<tbody><tr>
<td>AI Model</td>
<td>Gemini 2.5 Flash</td>
<td>Text, image, audio — single API. Streaming + function calling.</td>
</tr>
<tr>
<td>Backend</td>
<td>Node.js + Express</td>
<td>REST endpoints + WebSocket server. Session memory. Docker.</td>
</tr>
<tr>
<td>Frontend</td>
<td>React + Vite</td>
<td>WebSocket client. KaTeX maths rendering. Streaming display.</td>
</tr>
<tr>
<td>Cloud</td>
<td>Google Cloud Run</td>
<td>Serverless. Scale-to-zero. us-central1 free tier.</td>
</tr>
<tr>
<td>Security</td>
<td>Google Secret Manager</td>
<td>API key never in env vars or logs.</td>
</tr>
<tr>
<td>CI/CD</td>
<td>Google Cloud Build</td>
<td>Auto-triggered container build on deploy.</td>
</tr>
</tbody></table>
<h3>The AgentService and session memory</h3>
<p>The core of the backend is an <code>AgentService</code> class that owns all interaction with the Gemini API. Every request — text, image, or voice — flows through it.</p>
<pre><code class="language-javascript">class AgentService {
  async processMessage(sessionId, userInput, inputType) {
    const history = sessionStore.get(sessionId) || [];
    const systemPrompt = buildTutorSystemPrompt();
    const contents = buildContents(history, userInput, inputType);

    const stream = await geminiModel.generateContentStream({
      systemInstruction: systemPrompt,
      contents,
      tools: availableTools
    });

    let fullResponse = '';
    for await (const chunk of stream) {
      const token = chunk.text();
      this.emit('token', token);
      fullResponse += token;
    }

    sessionStore.save(sessionId, history, userInput, fullResponse);
  }
}
</code></pre>
<p>Session memory is held in-process with a 30-minute TTL. The full conversation history is passed back to Gemini on every request, giving the agent complete context of everything discussed.</p>
<h3>WebSocket for real-time streaming</h3>
<p>I migrated from REST to WebSocket mid-project when I realised REST couldn't support conversational continuity. Tokens stream from Gemini directly to the client as they are generated, and the persistent connection means a student's follow-up question mid-explanation arrives on the same channel with full session context preserved.</p>
<pre><code class="language-javascript">wss.on('connection', (ws) =&gt; {
  ws.on('message', async (data) =&gt; {
    const { type, text, sessionId, imageData } = JSON.parse(data);

    const agent = new AgentService();
    agent.on('token', (token) =&gt; {
      ws.send(JSON.stringify({ type: 'token', content: token }));
    });

    await agent.processMessage(sessionId, text, type, imageData);
    ws.send(JSON.stringify({ type: 'done' }));
  });
});
</code></pre>
<hr />
<h2>The feature that makes it a tutor: interruptible sessions</h2>
<p>The feature I spent the most time on is what I call <strong>interruptible tutoring</strong>. At any point during a step-by-step explanation, a student can send a follow-up question. The agent addresses it completely and then resumes the original explanation from the right place.</p>
<p>This works because of two things:</p>
<ul>
<li><p><strong>Session memory</strong> — the full conversation history, including the interrupted explanation, is passed back to Gemini with the follow-up. The model has complete context.</p>
</li>
<li><p><strong>System prompt design</strong> — the agent is explicitly instructed to address follow-up questions fully and then resume. The prompt frames this as a natural part of good tutoring.</p>
</li>
</ul>
<blockquote>
<p><strong>The key prompt insight:</strong> "Do not give the answer" produced stilted behaviour. "Your goal is for the student to reach understanding themselves — your job is to create the conditions for that discovery" produced natural, genuinely helpful tutoring. The framing of the goal, not just the constraint, was what mattered.</p>
</blockquote>
<hr />
<h2>Multimodal in practice: image and voice</h2>
<h3>Image input — handwritten homework</h3>
<p>The uploaded file is converted to base64 and passed as an inline data part in the Gemini API call. The agent states what it believes the problem to be before beginning the explanation — a verification step that lets students correct any misreading and eliminates a whole class of frustrating interactions.</p>
<h3>Voice input — spoken equations</h3>
<p>Voice input is sent as a base64-encoded audio file. Gemini handles transcription and interpretation natively, correctly processing spoken mathematical notation like "x squared plus five x" into x² + 5x, without any additional processing layer.</p>
<pre><code class="language-javascript">// Backend: /api/audio endpoint
app.post('/api/audio', async (req, res) =&gt; {
  const { audioData, sessionId } = req.body;   // base64 MP3

  const contents = [{
    role: "user",
    parts: [
      { inlineData: { mimeType: "audio/mp3", data: audioData } },
      { text: "Please transcribe and then help me solve this." }
    ]
  }];

  const agent = new AgentService();
  await agent.processWithContents(sessionId, contents, res);
});
</code></pre>
<hr />
<h2>Deploying to Google Cloud Run — staying within the free tier</h2>
<h3>Region selection is critical</h3>
<p>Cloud Run's always-free tier is only available in <code>us-central1</code>, <code>us-east1</code>, and <code>us-west1</code>. Deploying anywhere else — including European or Asian regions — is billable from the first request. The deploy script is locked to <code>us-central1</code>.</p>
<h3>Scale-to-zero is the most important cost control</h3>
<p>With <code>--min-instances=0</code>, the container shuts down completely when idle. You burn zero CPU-seconds during quiet periods.</p>
<pre><code class="language-bash">gcloud run deploy math-tutor-ai-backend \
  --region="us-central1" \           # free tier region
  --min-instances="0" \              # scale to zero = zero idle cost
  --max-instances="3" \              # hard cap prevents runaway billing
  --timeout="300s" \                 # supports WebSocket sessions
  --set-secrets="GEMINI_API_KEY=gemini-api-key:latest"  # Secret Manager
</code></pre>
<h3>Free tier numbers</h3>
<p>Cloud Run's always-free tier in <code>us-central1</code>: 180,000 vCPU-seconds, 360,000 GiB-seconds of memory, and 2 million requests per month — resetting every month. For a hackathon project with scale-to-zero, this is effectively unlimited.</p>
<blockquote>
<p><strong>Tip:</strong> Set a billing budget alert at $1 in the Google Cloud console before deploying. Takes 90 seconds and guarantees you'll never be surprised by a charge.</p>
</blockquote>
<hr />
<h2>KaTeX: making maths look like maths</h2>
<p>Gemini returns explanations with LaTeX notation naturally. The frontend renders it with <strong>KaTeX</strong>, a client-side LaTeX renderer. Equations appear properly typeset — not as ugly plain-text approximations. This detail signals to students that the tool takes mathematics seriously.</p>
<hr />
<h2>What I learned</h2>
<p><strong>Prompt engineering is product design.</strong> The most impactful work was writing the system prompt. The framing of the goal — not just the constraints — determines the entire character of the experience.</p>
<p><strong>Gemini's multimodal capability genuinely changes what you can build.</strong> Handling text, image, and audio in a single model call simplifies the architecture and reduces failure points in ways that matter for real-world applications.</p>
<p><strong>Serverless + AI APIs have eliminated the cost barrier for educational technology.</strong> A production-quality AI tutoring system can now be deployed and run at zero cost until it reaches meaningful scale. This is significant beyond this one project.</p>
<p><strong>Real-time streaming changes the psychology of the interaction.</strong> Watching an explanation build token by token feels like watching someone think. That quality — the sense of a live presence — matters enormously for an application whose value is that it feels like a real tutor.</p>
<hr />
<h2>What's next</h2>
<p>Persistent learning profiles. Adaptive difficulty. Expansion into physics and chemistry. Interactive graphing. A lightweight offline mode for low-bandwidth environments.</p>
<p>The goal has not changed from the first line of code to the last: <strong>every student, everywhere, deserves a patient and knowledgeable tutor.</strong></p>
<hr />
<p><em>This post was created for the purposes of entering the</em> <em><strong>Gemini Live Agent Challenge</strong></em> <em>hackathon.</em></p>
<p><strong>#GeminiLiveAgentChallenge</strong> #GoogleCloud #GeminiAI #BuildWithGoogle #EdTech #AI #NodeJS #React #CloudRun</p>
<hr />
<p><strong>Links</strong></p>
<ul>
<li><p>Live app: <a href="https://ai-agent-frontend-1087118236338.us-central1.run.app">https://ai-agent-frontend-1087118236338.us-central1.run.app</a></p>
</li>
<li><p>Backend repo: <a href="https://github.com/tay4real/ai-agent-backend">https://github.com/tay4real/ai-agent-backend</a></p>
</li>
<li><p>Frontend repo: <a href="https://github.com/tay4real/ai-agent-frontend">https://github.com/tay4real/ai-agent-frontend</a></p>
</li>
</ul>
]]></content:encoded></item><item><title><![CDATA[The Ultimate 2023 Guide to Installing React Native: 10 Simple Steps to Success]]></title><description><![CDATA[Introduction to React Native
React Native, for the uninitiated, is a game-changing framework developed by Facebook. It enables developers to create natively-rendered mobile applications for both iOS and Android platforms using a single JavaScript cod...]]></description><link>https://ademuyiwaotubusin.hashnode.dev/the-ultimate-2023-guide-to-installing-react-native-10-simple-steps-to-success</link><guid isPermaLink="true">https://ademuyiwaotubusin.hashnode.dev/the-ultimate-2023-guide-to-installing-react-native-10-simple-steps-to-success</guid><category><![CDATA[mobile app development]]></category><category><![CDATA[React Native]]></category><dc:creator><![CDATA[Ademuyiwa Otubusin]]></dc:creator><pubDate>Mon, 21 Aug 2023 13:14:13 GMT</pubDate><enclosure url="https://cdn.hashnode.com/res/hashnode/image/upload/v1692623430740/00d3414b-3de0-4b80-8e82-ff932eb15987.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h3 id="heading-introduction-to-react-native"><strong>Introduction to React Native</strong></h3>
<p>React Native, for the uninitiated, is a game-changing framework developed by Facebook. It enables developers to create natively-rendered mobile applications for both iOS and Android platforms using a single JavaScript codebase.</p>
<h4 id="heading-what-is-react-native"><strong>What is React Native?</strong></h4>
<p>In the simplest terms, React Native bridges the gap between web development and native app creation. It leverages the same design as the popular React framework but directs these components to native end-points.</p>
<h4 id="heading-why-use-react-native"><strong>Why Use React Native?</strong></h4>
<p>Opting for React Native isn't just a fad; it's a strategic move. The framework saves significant development time, reduces cost, and offers high performance. Moreover, its hot-reloading feature ensures that you see the results of the latest change immediately, without losing your current application state.</p>
<hr />
<h3 id="heading-installation-guide-for-react-native"><strong>Installation Guide For React Native</strong></h3>
<p>Embarking on your React Native journey involves setting up your environment correctly. Let's dive in!</p>
<h4 id="heading-prerequisites"><strong>Prerequisites</strong></h4>
<p>Before getting hands-on with React Native, you'll need:</p>
<ul>
<li><p><strong>Operating System</strong>: Windows, macOS, or Linux.</p>
</li>
<li><p><strong>Code Editor</strong>: Visual Studio Code, Atom, etc.</p>
</li>
<li><p><strong>Terminal</strong>: Command Prompt, Terminal, or iTerm.</p>
</li>
</ul>
<h4 id="heading-setting-up-your-environment"><strong>Setting Up Your Environment</strong></h4>
<ol>
<li><p>Ensure that you have a JDK (Java Development Kit) installed. Check by typing <code>java -version</code> in your terminal.</p>
</li>
<li><p>Install Android Studio for Android development. This will provide you with the Android SDK and emulators.</p>
</li>
<li><p>For iOS development, you'll need a Mac with Xcode installed.</p>
</li>
</ol>
<h4 id="heading-installing-nodejs-and-the-react-native-cli"><strong>Installing Node.js and the React Native CLI</strong></h4>
<p>Node.js is pivotal to React Native development. To install:</p>
<ol>
<li><p>Head to <a target="_blank" href="https://nodejs.org/"><strong>Node.js official website</strong></a> and download the recommended version.</p>
</li>
<li><p>Once installed, open your terminal and type <code>npm install -g react-native-cli</code>. This installs the React Native command-line interface.</p>
</li>
</ol>
<h4 id="heading-android-development-environment"><strong>Android Development Environment</strong></h4>
<p>Ensure that you have the Android Studio SDK, an Android Device or emulator, and the necessary build tools.</p>
<h4 id="heading-ios-development-environment"><strong>iOS Development Environment</strong></h4>
<p>For iOS, you'd need a Mac with Xcode and an iOS simulator or device. Ensure that you've also set up CocoaPods, an application-level dependency manager for Swift and Objective-C.</p>
<h4 id="heading-creating-a-new-react-native-project"><strong>Creating a New React Native Project</strong></h4>
<p>Once everything's set up, create a new project with <code>react-native init YourProjectName</code>.</p>
<hr />
<h3 id="heading-building-with-react-native"><strong>Building with React Native</strong></h3>
<h4 id="heading-components-and-syntax"><strong>Components and Syntax</strong></h4>
<p>Building in React Native is all about components. These reusable pieces define how UI renders and functions.</p>
<h4 id="heading-state-management-and-redux"><strong>State Management and Redux</strong></h4>
<p>Redux plays a vital role in state management in React Native applications, ensuring smooth data flow throughout the app.</p>
<hr />
<h3 id="heading-testing-your-react-native-app"><strong>Testing Your React Native App</strong></h3>
<h4 id="heading-tools-for-testing"><strong>Tools for Testing</strong></h4>
<p>There's a plethora of tools out there, like Jest and Mocha. Choosing one depends on your project's needs.</p>
<h4 id="heading-writing-and-running-tests"><strong>Writing and Running Tests</strong></h4>
<p>It's crucial to maintain a testing regimen, ensuring your app's robustness and reliability.</p>
<hr />
<h3 id="heading-deployment-and-distribution"><strong>Deployment and Distribution</strong></h3>
<h4 id="heading-deploying-to-app-stores"><strong>Deploying to App Stores</strong></h4>
<p>Deploying your app requires adhering to guidelines set by Apple and Google. Always keep them in mind!</p>
<h4 id="heading-over-the-air-updates-with-codepush"><strong>Over-the-Air Updates with CodePush</strong></h4>
<p>Microsoft's CodePush allows developers to push code updates to devices instantly.</p>
<hr />
<h3 id="heading-best-practices"><strong>Best Practices</strong></h3>
<h4 id="heading-performance-optimization"><strong>Performance Optimization</strong></h4>
<p>To ensure your app runs buttery smooth, consider techniques like lazy loading.</p>
<h4 id="heading-common-pitfalls-to-avoid"><strong>Common Pitfalls to Avoid</strong></h4>
<p>Steer clear from making common mistakes like neglecting platform-specific design guidelines.</p>
]]></content:encoded></item></channel></rss>