Atlassian interview questions & answers

20 real Atlassian interview questions with full model answers — System design, Behavioral, Technical, Coding. Drawn from the same verified bank ChannelPulse drills from (65 Atlassian questions in total).

BehavioralEasyAtlassian

1. Tell me about a time when you had to collaborate with a team to solve a challenging problem.

The full question

Tell me about a time when you had to collaborate with a team to solve a challenging problem. What was your role, and what was the outcome?

Model answer

Situation In my role as a software engineer at a mid-sized tech company, our team was tasked with developing a real-time data analytics platform for a major client. This project was critical as it was intended to provide advanced insights and visualizations, which were key selling points for our client. The challenge was integrating a third-party data visualization library with our custom backend solution, which required seamless collaboration across multiple teams, including frontend and backend developers, UX designers, and data scientists.

Task I was responsible for ensuring that the integration was smooth and that the platform met the client's expectations for real-time data processing and visualization. The key constraint was the tight deadline, as the client needed the platform ready for an upcoming product launch.

Action

  • I initiated a brainstorming session with all relevant stakeholders, including developers, UX designers, and data scientists, to explore different integration approaches and identify potential roadblocks early on.
  • I facilitated open communication by setting up a shared workspace where team members could continuously update progress and share insights, ensuring everyone was aligned and informed.
  • I coordinated with the UX team to ensure that the user interface was intuitive and aligned with the client's branding, while also supporting the technical requirements of the data visualization library.
  • I led the technical integration efforts, focusing on optimizing data flow between the backend and the visualization library to ensure real-time performance.
  • I organized regular check-ins and feedback sessions with the client to ensure that their expectations were being met and to make adjustments as needed based on their input.

Result We successfully delivered the real-time data analytics platform within the given timeline. The client was delighted with the platform’s user-friendly interface and the advanced visualizations it provided. This project not only met the client's needs but also strengthened our team's ability to collaborate across different disciplines. Reflecting on this experience, I learned the importance of effective communication and collaboration in overcoming complex challenges, which are just as crucial as technical skills in software development.

BehavioralMediumAtlassian

2. Describe a situation where you had to prioritize multiple tasks with tight deadlines.

The full question

Describe a situation where you had to prioritize multiple tasks with tight deadlines. How did you manage your time and resources?

Model answer

Situation In my role as a software developer at a mid-sized tech company, I encountered a challenging situation where I had to manage multiple high-priority tasks with tight deadlines. Our team was in the midst of delivering a critical update for a key client, which coincided with the final stages of another project involving a new feature launch. Both projects were crucial for the company’s strategic goals, and failure to deliver on time could have resulted in significant client dissatisfaction and potential revenue loss.

Task I was responsible for ensuring the timely completion of the backend components for both projects. The key constraint was the limited time available, as both deadlines were set within a week of each other, and the tasks required significant coordination with other team members.

Action

  • I began by reassessing the priorities of each task, focusing on those with the highest impact and urgency. This involved breaking down each project into smaller, manageable tasks and identifying dependencies.
  • I coordinated with my team to redistribute the workload effectively. We identified areas where we could leverage each other's strengths and sought additional resources from other teams to help with specific components.
  • To maximize efficiency, I streamlined my work processes and extended my work hours temporarily. I also utilized productivity tools to keep track of progress and ensure that I was meeting interim milestones.
  • I maintained regular communication with stakeholders, providing updates on our progress and any potential risks to the timeline. This transparency helped manage expectations and allowed for quick adjustments when necessary.
  • Additionally, I implemented a risk assessment strategy to anticipate potential roadblocks and devised contingency plans to address them proactively.

Result Through these efforts, we successfully delivered the critical update on time for the client's promotional event, which was well-received and resulted in positive feedback. The new feature launch was completed shortly after, with only a minor delay that did not impact the overall project timeline significantly. This experience taught me the importance of prioritization, effective communication, and the ability to adapt quickly under pressure, all of which have been invaluable in subsequent projects.

BehavioralMediumAtlassian

3. Can you share an experience where you had to advocate for a technical decision that was initially met with resistance?

The full question

Can you share an experience where you had to advocate for a technical decision that was initially met with resistance? How did you handle it?

Model answer

Situation

In my previous role as a software developer at a mid-sized tech company, our team was tasked with enhancing the scalability of our primary web application. During a planning session, the majority of the team favored implementing a microservices architecture to achieve this goal. However, I had significant concerns about the complexity and potential over-engineering of this solution given our current scale and resources.

Task

My task was to advocate for a more incremental approach that involved optimizing our existing monolithic architecture. I needed to convince the team that this approach could meet our scalability needs without the risks associated with a complete architectural overhaul.

Action

  • I began by conducting a thorough analysis of our current system's performance bottlenecks and identified specific areas where targeted optimizations could yield significant improvements.
  • I prepared a detailed presentation outlining the benefits of optimizing the existing architecture, including reduced implementation time, lower risk, and cost-effectiveness. I also highlighted potential pitfalls of transitioning to microservices prematurely, such as increased complexity and maintenance overhead.
  • During the next team meeting, I presented my findings and proposed a phased approach. I suggested starting with optimizations and then gradually moving to microservices for specific components if needed.
  • I encouraged open discussion and invited feedback from the team. I made sure to address concerns and questions, emphasizing that my proposal was not against innovation but rather a strategic step to ensure stability and efficiency.
  • To bolster my argument, I shared case studies from similar companies that successfully scaled their systems using a similar approach.

Result

My manager and several team members appreciated the comprehensive analysis and the pragmatic approach I proposed. After further discussions, the team agreed to adopt the incremental optimization strategy. This decision led to a 30% improvement in system performance within three months, with minimal disruption to ongoing operations. This experience taught me the value of thorough preparation, clear communication, and the importance of considering both short-term and long-term impacts when advocating for technical decisions.

BehavioralMediumAtlassianMachine Learning EngineerOnsite

4. You are in a management/values interview.

The full question

You are in a management/values interview. Prepare behavioral answers aligned to the company’s core values.

Typical questions in this round include:

  • Describe a time you pushed back on a decision or requirement.
  • Describe a time you drove change / improvement.
  • Describe a time you helped other people improve (mentoring/coaching).
  • Describe a time you worked through ambiguity.
  • Describe a mistake you made and how you corrected course.

Interviewers may probe with follow-ups if they haven’t yet heard the signal they need (scope, ownership, impact, conflict handling, learning).

Model answer

Situation

In my previous role as a product manager at a mid-sized tech company, I was responsible for overseeing the development of a new feature for our flagship product. The project was crucial as it aimed to enhance user engagement, which was a key metric for our business growth. However, during the planning phase, the engineering team proposed a technical approach that I believed would not scale well with our projected user growth.

Task

My task was to ensure that the project not only met immediate needs but was also sustainable in the long term. This required me to push back on the proposed approach and advocate for a more scalable solution, despite the initial resistance from the engineering team.

Action

  • I first gathered data on user growth projections and the potential load on the system, presenting a clear case for why the current approach might lead to performance issues.
  • I scheduled a meeting with the engineering lead to discuss my concerns, ensuring I approached the conversation with respect and openness to their perspective.
  • During the meeting, I highlighted the importance of aligning our technical decisions with our long-term business goals, emphasizing the value of investing in a scalable solution.
  • I proposed an alternative approach that involved using a microservices architecture, which would allow us to handle increased traffic more effectively.
  • To gain buy-in, I collaborated with a senior engineer to create a prototype that demonstrated the feasibility and benefits of the proposed solution.
  • Throughout the process, I maintained clear and open communication with all stakeholders, ensuring everyone was aligned and informed about the potential impacts and benefits.

Result

As a result of these efforts, the team agreed to adopt the microservices approach. This decision not only improved the system's scalability but also enhanced its reliability and performance. The feature launch was successful, leading to a 20% increase in user engagement within the first three months. Reflecting on this experience, I learned the importance of balancing immediate project needs with long-term strategic goals and the value of constructive dialogue in driving positive change.

CodingEasyAtlassian

5. Given an array of integers, return indices of the two numbers such that they add up to a specific target.

Model answer

function twoSum(nums, target) {
  // Create a map to store the difference and its index
  const numMap = new Map();

  // Iterate over the array
  for (let i = 0; i < nums.length; i++) {
    // Calculate the difference needed to reach the target
    const complement = target - nums[i];

    // Check if the complement is already in the map
    if (numMap.has(complement)) {
      // If found, return the indices of the two numbers
      return [numMap.get(complement), i];
    }

    // Otherwise, add the current number and its index to the map
    numMap.set(nums[i], i);
  }

  // If no solution is found, return an empty array
  return [];
}

// Example usage:
// const indices = twoSum([2, 7, 11, 15], 9);
// console.log(indices); // Output: [0, 1]
  • Approach:
  • Use a hash map to store each number and its index as you iterate through the array.
  • For each number, calculate the complement (target - current number).
  • Check if the complement exists in the map; if it does, return the indices.
  • If not, add the current number and its index to the map.
  • This ensures each number is processed in constant time, making the solution efficient.
  • Complexity:
  • Time: O(n), where n is the number of elements in the array, since we traverse the list once.
  • Space: O(n), for storing elements in the hash map.
CodingEasyAtlassian

6. Given an array of integers, find the maximum sum of a non-empty subarray.

Model answer

function maxSubArray(nums) {
    // Initialize variables to store the maximum sum of subarray found so far
    // and the current sum of the subarray being considered
    let maxSum = nums[0];
    let currentSum = nums[0];

    // Iterate through the array starting from the second element
    for (let i = 1; i < nums.length; i++) {
        // Update the current sum by including the current element
        // or start a new subarray from the current element if it gives a higher sum
        currentSum = Math.max(nums[i], currentSum + nums[i]);

        // Update the maximum sum found so far
        maxSum = Math.max(maxSum, currentSum);
    }

    return maxSum;
}

// Example usage:
console.log(maxSubArray([-2, 1, -3, 4, -1, 2, 1, -5, 4])); // Output: 6
  • Approach: This solution uses Kadane's Algorithm, which efficiently finds the maximum sum of a contiguous subarray in linear time.
  • Initialization: Start with the first element as both the current sum and the maximum sum.
  • Iteration: For each element, decide whether to add it to the current subarray or start a new subarray. Update the maximum sum accordingly.
  • Complexity:
  • Time: O(n), where n is the number of elements in the array, as we iterate through the array once.
  • Space: O(1), as we use a constant amount of extra space.
CodingEasyAtlassian

7. Given an array of integers, write a function to find the maximum sum of any contiguous subarray of the array.

The full question

Given an array of integers, write a function to find the maximum sum of any contiguous subarray of the array. Return 0 if the array is empty.

Model answer

function maxSubArray(nums) {
    if (nums.length === 0) return 0;

    let maxSum = nums[0]; // Initialize maxSum to the first element
    let currentSum = nums[0]; // Initialize currentSum to the first element

    for (let i = 1; i < nums.length; i++) {
        // Choose the larger between the current element and the sum of current element with currentSum
        currentSum = Math.max(nums[i], currentSum + nums[i]);
        // Update maxSum if currentSum is greater
        maxSum = Math.max(maxSum, currentSum);
    }

    return maxSum;
}

// Example usage:
console.log(maxSubArray([-2, 1, -3, 4, -1, 2, 1, -5, 4])); // Output: 6
  • Approach: This solution uses Kadane’s algorithm, which efficiently finds the maximum sum of a contiguous subarray in linear time.
  • Initialization: Start with the first element as both maxSum and currentSum.
  • Iteration: For each element, decide whether to add it to the current subarray (currentSum + nums[i]) or start a new subarray with the current element (nums[i]).
  • Update: Continuously update maxSum with the maximum value found.

Complexity:

  • Time: O(n), where n is the number of elements in the array, since we iterate through the array once.
  • Space: O(1), as we use a constant amount of extra space.
CodingEasyAtlassian

8. Reverse a given string.

Model answer

function reverseString(str) {
  // Convert the string to an array of characters
  let charArray = str.split('');
  
  // Initialize two pointers
  let left = 0;
  let right = charArray.length - 1;
  
  // Swap characters until the pointers meet in the middle
  while (left < right) {
    // Swap the characters at left and right indices
    let temp = charArray[left];
    charArray[left] = charArray[right];
    charArray[right] = temp;
    
    // Move the pointers towards the center
    left++;
    right--;
  }
  
  // Join the array back into a string and return
  return charArray.join('');
}

// Example usage:
console.log(reverseString("hello")); // Output: "olleh"
  • Approach:
  • Convert the string into an array of characters to allow in-place modifications.
  • Use two pointers, one starting at the beginning (left) and the other at the end (right).
  • Swap the characters at these pointers and move them towards the center.
  • Continue swapping until the pointers meet or cross each other.
  • Join the modified array back into a string and return it.
  • Complexity:
  • Time: O(n), where n is the length of the string, as each character is processed once.
  • Space: O(n), due to the array created to hold the characters of the string.
Product & growthEasyAtlassianProduct Manager

9. What is your favorite Atlassian product and why?

The full question

What is your favorite Atlassian product and why? How would you improve it?

Model answer

Favorite Product: My favorite Atlassian product is Trello due to its intuitive design and flexibility in project management.

Why: Trello's visual boards and card system make task management simple and accessible for teams of all sizes.

Improvement: To enhance Trello, I would focus on improving its reporting capabilities.

Clarify & scope: The goal is to provide better insights into project progress. Assume current reporting tools are basic and require manual data extraction.

User segments & pain points: Target project managers who need detailed reports to track team performance.

Goals & success metrics: The North Star metric is the frequency of report generation. Guardrails include user satisfaction and ease of use.

Solutions:

  1. Advanced Reporting Dashboard: Offer customizable dashboards with real-time data visualization.
  2. Automated Report Generation: Enable scheduled reports sent directly to stakeholders.
  3. Integration with BI Tools: Allow export to business intelligence tools for deeper analysis.

Recommendation: Implement an advanced reporting dashboard for immediate impact.

Prioritization & trade-offs: Prioritize the dashboard for its high impact and moderate effort.

MVP, measurement & rollout: Launch a beta version of the dashboard. Measure usage and gather feedback to refine the feature.

Product & growthMediumAtlassianProduct Manager

10. How would you improve Jira's onboarding experience for new users?

Model answer

Clarify & scope: The goal is to enhance Jira's onboarding experience to increase user retention and satisfaction. Assume the current onboarding process is lengthy and complex, leading to user drop-off.

User segments & pain points: Focus on new users who are unfamiliar with project management tools. Their pain points include confusion with the interface and difficulty understanding the tool's value.

Goals & success metrics: The North Star metric is the completion rate of the onboarding process. Guardrails include user satisfaction scores and time taken to complete onboarding.

Solutions:

  1. Interactive Tutorials: Implement step-by-step guides that walk users through key features.
  2. Personalized Onboarding: Tailor the onboarding flow based on user roles and project types.
  3. Gamification: Use achievements and rewards to motivate users to complete onboarding.

Recommendation: Implement interactive tutorials as they provide immediate guidance and can be adapted easily.

graph LR
A[New User] --> B[Interactive Tutorial]
B --> C[Feature Walkthrough]
C --> D[Onboarding Completion]
Diagram

Prioritization & trade-offs: Using RICE, prioritize interactive tutorials for high reach and impact with moderate effort.

MVP, measurement & rollout: Launch a basic version of interactive tutorials. Measure completion rates and gather feedback for iteration.

Product & growthMediumAtlassianProduct Manager

11. Design a feature for Confluence that enhances team collaboration on documents.

Model answer

Clarify & scope: The goal is to enhance team collaboration within Confluence, focusing on document co-editing. Assume current collaboration tools are limited in real-time interaction.

User segments & pain points: Target remote teams who struggle with asynchronous communication and document versioning issues.

Goals & success metrics: The North Star metric is the frequency of document co-editing sessions. Guardrails include user engagement rates and feedback scores.

Solutions:

  1. Real-time Co-editing: Allow multiple users to edit documents simultaneously with change tracking.
  2. Comment Threads: Enable threaded discussions directly on document sections.
  3. Version Control: Simplify version management with automatic saving and rollback options.

Recommendation: Implement real-time co-editing to address immediate collaboration needs.

graph LR
A[User 1] --> B[Document]
A --> C[Edit]
B --> D[Real-time Changes]
C --> D
Diagram

Prioritization & trade-offs: Prioritize real-time co-editing for its high impact on collaboration, despite the technical effort required.

MVP, measurement & rollout: Develop a prototype for real-time co-editing. Measure engagement and collect user feedback for improvements.

Product & growthMediumAtlassianProduct Manager

12. How would you improve the integration of Atlassian tools with third-party applications?

Model answer

Clarify & scope: The goal is to enhance integration capabilities between Atlassian tools and third-party applications. Assume the current integration process is cumbersome and limited.

User segments & pain points: Focus on developers and IT administrators who require seamless tool interoperability.

Goals & success metrics: The North Star metric is the number of successful third-party integrations. Guardrails include integration setup time and user satisfaction.

Solutions:

  1. Unified API Platform: Develop a comprehensive API platform for easier integration.
  2. Pre-built Connectors: Offer a library of connectors for popular third-party apps.
  3. Integration Marketplace: Create a marketplace for community-built integrations.

Recommendation: Develop pre-built connectors to quickly address common integration needs.

Prioritization & trade-offs: Prioritize pre-built connectors for their high reach and moderate development effort.

MVP, measurement & rollout: Launch a set of connectors for top-requested apps. Measure integration success rates and gather feedback for expansion.

System designEasyAtlassian

13. Design a simple task management system where users can create, update, and delete tasks.

The full question

Design a simple task management system where users can create, update, and delete tasks. What components would you include?

Model answer

1. Requirements & scale

Functional Requirements:

  • Users can create, update, and delete tasks.
  • Users can view a list of tasks.
  • Tasks should have attributes like title, description, due date, and status.

Non-Functional Requirements:

  • High availability and reliability.
  • Low latency for task operations.
  • Scalability to support a growing number of users.

Estimates:

  • Assume 100,000 users, each making an average of 10 task operations per day.
  • This results in approximately 1,000,000 operations per day.
  • QPS (Queries Per Second) = 1,000,000 / (24 60 60) ≈ 11.6.
  • Storage: Assume each task is about 1 KB. With 1,000,000 tasks, storage needed is approximately 1 GB.

2. High-level architecture

flowchart TD
    subgraph Client
        A[User Interface]
    end

    subgraph Edge/CDN
        B[CDN]
    end

    subgraph Load Balancer
        C[Load Balancer]
    end

    subgraph API / Services
        D[Task Service]
    end

    subgraph Cache
        E[Redis Cache]
    end

    subgraph Datastores
        F[SQL Database]
    end

    A -->|HTTP Requests| B
    B -->|Forward Requests| C
    C -->|API Calls| D
    D -->|Read/Write| E
    D -->|Read/Write| F
    E -->|Cache Miss| F
Diagram

3. API design

  • POST /tasks: Create a new task.
  • GET /tasks: Retrieve a list of tasks.
  • GET /tasks/{id}: Retrieve a specific task by ID.
  • PUT /tasks/{id}: Update a task by ID.
  • DELETE /tasks/{id}: Delete a task by ID.

4. Data model & storage

Datastore Choice:

  • Use a SQL database (e.g., PostgreSQL) for ACID compliance and structured data storage.

Key Tables:

  • Tasks Table:
  • id (Primary Key)
  • title (VARCHAR)
  • description (TEXT)
  • due_date (DATE)
  • status (ENUM: 'pending', 'completed')
  • created_at (TIMESTAMP)
  • updated_at (TIMESTAMP)

Partitioning:

  • Use id as the primary key for partitioning and indexing to ensure quick lookups and updates.

5. Deep dive

The core functionality of the task management system revolves around CRUD operations. To ensure efficient data retrieval and updates, caching can be employed. Here's a sequence diagram illustrating the flow for retrieving a task:

sequenceDiagram
    participant User
    participant CDN
    participant LoadBalancer
    participant TaskService
    participant Cache
    participant Database

    User->>CDN: GET /tasks/{id}
    CDN->>LoadBalancer: Forward Request
    LoadBalancer->>TaskService: API Call
    TaskService->>Cache: Check Cache for Task
    alt Cache Hit
        Cache-->>TaskService: Return Task
    else Cache Miss
        TaskService->>Database: Query Task
        Database-->>TaskService: Return Task
        TaskService->>Cache: Update Cache
    end
    TaskService->>User: Return Task
Diagram

6. Scale, bottlenecks & trade-offs

Scalability:

  • Replication: Use database replication to ensure high availability and distribute read loads.
  • Sharding: Implement sharding if the dataset grows significantly, using the id field for shard keys.

Caching:

  • Use Redis to cache frequently accessed tasks to reduce database load and improve response times.

Bottlenecks:

  • The database could become a bottleneck under heavy load. Mitigate this by optimizing queries and using read replicas.
  • The cache layer could become a single point of failure. Ensure redundancy and failover mechanisms are in place.

Trade-offs:

  • Consistency vs. Availability: Opt for eventual consistency in caching to improve availability.
  • SQL vs. NoSQL: SQL is chosen for its ACID properties, but NoSQL could be considered if flexibility and horizontal scaling become priorities.

By focusing on these components and considerations, the task management system can efficiently handle user operations while being scalable and reliable.

System designMediumAtlassian

14. Design a data structure that supports the following operations: insert, delete, get_random_element.

The full question

Design a data structure that supports the following operations: insert, delete, get_random_element. All operations should be done in average O(1) time.

Model answer

1. Requirements & scale

Functional Requirements:

  • Insert: Add an element to the data structure.
  • Delete: Remove an element from the data structure.
  • Get Random Element: Retrieve a random element from the data structure.

Non-Functional Requirements:

  • Efficiency: All operations should be performed in average O(1) time.
  • Scalability: The data structure should handle a large number of elements efficiently.

Estimates:

  • Operations per second (QPS): Assume 10,000 operations per second.
  • Storage: If each element is an integer (4 bytes), storing 1 million elements requires about 4 MB.
  • Bandwidth: Minimal, as operations are primarily in-memory.

2. High-level architecture

flowchart TD
    subgraph Client
        A[Client]
    end

    subgraph API / Services
        B[Data Structure Service]
    end

    A -->|Insert/Delete/Get Random| B
Diagram

3. API design

  • POST /insert: Add an element to the data structure.
  • DELETE /delete: Remove an element from the data structure.
  • GET /get_random_element: Retrieve a random element from the data structure.

4. Data model & storage

We will use a combination of a dynamic array (or list) and a hash map (or dictionary) to achieve average O(1) time complexity for all operations.

  • Array/List: Used to store elements for quick access and random selection.
  • Hash Map/Dictionary: Maps elements to their indices in the array for quick lookup and deletion.

Data Structure Design:

  • Array/List: elements[]
  • Hash Map/Dictionary: elementIndexMap{element: index}

5. Deep dive

To achieve O(1) operations, we use the following approach:

  • Insert: Add the element to the end of the array and update the hash map with the element and its index.
  • Delete: Find the element's index using the hash map. Swap the element with the last element in the array, update the hash map for the swapped element, and then remove the last element from the array.
  • Get Random Element: Use a random number generator to pick an index within the bounds of the array and return the element at that index.
sequenceDiagram
    participant Client
    participant DataStructureService
    Client->>DataStructureService: POST /insert(element)
    DataStructureService-->>Client: Element inserted

    Client->>DataStructureService: DELETE /delete(element)
    DataStructureService-->>Client: Element deleted

    Client->>DataStructureService: GET /get_random_element
    DataStructureService-->>Client: Random element
Diagram

6. Scale, bottlenecks & trade-offs

Scalability: The data structure is designed to handle a large number of elements efficiently due to its O(1) operations.

Bottlenecks: The primary bottleneck could be memory usage if the data structure grows significantly large. However, since operations are in-memory, they are fast.

Trade-offs:

  • Consistency vs. Availability: The design is inherently consistent as all operations are atomic and in-memory.
  • Space vs. Time Complexity: The use of both an array and a hash map increases space usage but ensures time efficiency.
  • Randomness: The randomness of get_random_element is dependent on the quality of the random number generator.

Replication and Sharding: If needed, the data structure can be replicated across multiple servers for high availability. Sharding can be applied by partitioning elements based on a hash function, but this would complicate the random element retrieval process, as it would require coordination across shards.

System designMediumAtlassianSoftware EngineerOnsite

15. Design a service that crawls images starting from a set of root URLs.

The full question

Design a service that crawls images starting from a set of root URLs.

Requirements:

  • Input: one or more root URLs.
  • Crawl pages, discover links, and download image resources.
  • Support unlimited number of root URLs and unlimited crawl depth.
  • Must handle failures (network errors, timeouts, crashes) and avoid re-crawling the same URL excessively.
  • Discuss storage for downloaded images and metadata.

Deliverables:

  • High-level architecture (components, data flow).
  • Queue/scheduler design and politeness (per-host rate limiting).
  • Deduplication strategy.
  • DB schema for crawl state and results.
  • Failure/retry model and monitoring.

Model answer

1. Requirements & scale

Functional Requirements:

  • Accept one or more root URLs as input.
  • Crawl web pages starting from these URLs, discover links, and download image resources.
  • Support unlimited number of root URLs and unlimited crawl depth.
  • Handle failures such as network errors, timeouts, and crashes.
  • Avoid re-crawling the same URL excessively.

Non-Functional Requirements:

  • High availability and reliability.
  • Efficient storage and retrieval of images and metadata.
  • Rate limiting to avoid overloading web servers.

Estimates:

  • Assume we start with 100 root URLs, each leading to 1000 pages with 10 images on average.
  • Total images: 100 1000 10 = 1,000,000 images.
  • If each image is approximately 100 KB, total storage required: 1,000,000 * 100 KB = 100 GB.
  • Assuming each page takes 1 second to process, the system should handle 1000 QPS for crawling.

2. High-level architecture

flowchart TD
    subgraph Client
        A[User Interface]
    end

    subgraph Edge/CDN
        B[CDN]
    end

    subgraph Load Balancer
        C[Load Balancer]
    end

    subgraph API / Services
        D[API Gateway]
        E[Crawler Service]
        F[Scheduler]
    end

    subgraph Cache
        G[URL Cache]
    end

    subgraph Datastores
        H["Image Storage (Blob)"]
        I["Metadata DB (SQL)"]
    end

    subgraph Message Queue
        J[URL Queue]
    end

    subgraph Workers
        K[Crawl Workers]
    end

    A --> B --> C --> D --> E
    E --> F
    F --> J
    J --> K
    K --> G
    K --> H
    K --> I
    G --> K
Diagram

3. API design

  • POST /crawl: Start a new crawl with one or more root URLs.
  • GET /status/{crawlId}: Retrieve the status of a specific crawl.
  • GET /images/{imageId}: Fetch a downloaded image.
  • GET /metadata/{url}: Retrieve metadata for a specific URL.

4. Data model & storage

Datastores:

  • Blob Storage: For storing downloaded images. Chosen for scalability and cost-effectiveness.
  • SQL Database: For storing metadata and crawl state. Chosen for structured data and transactional support.

Key Tables:

  • Images: image_id, url, storage_path, download_timestamp.
  • CrawlState: url, status, last_crawled, retry_count.
  • Metadata: url, title, description, keywords.

Partition Key:

  • Use url as the partition key for deduplication and efficient retrieval.

5. Deep dive

The core of this design is the crawling and deduplication process. The system uses a URL queue to manage the crawl frontier. Each URL is processed by a crawl worker, which fetches the page, extracts image URLs, and stores them in the URL cache to avoid re-crawling.

sequenceDiagram
    participant User
    participant API
    participant Scheduler
    participant Queue
    participant Worker
    participant Cache
    participant Storage

    User->>API: POST /crawl
    API->>Scheduler: Initiate crawl
    Scheduler->>Queue: Add root URLs
    loop Crawl URLs
        Queue->>Worker: Fetch URL
        Worker->>Cache: Check URL
        alt URL not in Cache
            Worker->>Cache: Add URL
            Worker->>Storage: Download and store image
            Worker->>Queue: Add discovered URLs
        else URL in Cache
            Worker->>Queue: Skip URL
        end
    end
Diagram

6. Scale, bottlenecks & trade-offs

Scaling:

  • Replication: Use database replication for high availability and disaster recovery.
  • Sharding: Shard the URL queue and metadata database based on URL hash to distribute load.

Bottlenecks:

  • Network Bandwidth: Ensure sufficient bandwidth for downloading images.
  • Queue Size: Monitor and scale the message queue to handle high volumes of URLs.

Trade-offs:

  • Consistency vs. Availability: Opt for eventual consistency in the URL cache to improve availability.
  • Politeness: Implement per-host rate limiting to avoid overloading web servers, which may slow down the crawl process.
  • Failure Handling: Use exponential backoff for retries to handle transient network failures gracefully.

By addressing these aspects, the system can efficiently crawl and store images while managing resources and handling failures effectively.

System designMediumAtlassian

16. Design a notification system for an Atlassian product.

The full question

Design a notification system for an Atlassian product. How would you ensure timely delivery and reliability?

Model answer

1. Requirements & scale

Functional Requirements:

  • Deliver notifications to users in real-time.
  • Support multiple notification types (e.g., email, push, in-app).
  • Ensure message ordering and delivery guarantees.
  • Allow users to manage notification preferences.

Non-Functional Requirements:

  • High availability and reliability.
  • Low latency for real-time delivery.
  • Scalability to handle millions of users and notifications.
  • Fault tolerance and resilience.

Back-of-the-envelope Estimates:

  • Assume 10 million active users, each receiving an average of 10 notifications per day.
  • Total notifications per day: 100 million.
  • Peak QPS (Queries Per Second): Assume peak load is 10x average, leading to approximately 11,574 QPS.
  • Storage: If each notification is 1 KB, daily storage requirement is 100 GB.
  • Bandwidth: Assuming 1 KB per notification, peak bandwidth requirement is approximately 11.3 MB/s.

2. High-level architecture

flowchart TD
  subgraph Client
    A[User Devices]
  end
  
  subgraph Edge/CDN
    B[CDN]
  end
  
  subgraph Load Balancer
    C[Load Balancer]
  end
  
  subgraph API / Services
    D[Notification Service]
    E[User Preferences Service]
  end
  
  subgraph Cache
    F[Redis Cache]
  end
  
  subgraph Datastores
    G[SQL DB]
    H["Object Storage (S3)"]
  end
  
  subgraph Message Queue
    I[Kafka]
  end
  
  subgraph Workers
    J[Notification Workers]
  end
  
  A -->|Request Notification| B
  B --> C
  C --> D
  D -->|Check Preferences| E
  E -->|Fetch Preferences| F
  D -->|Store Notification| G
  D -->|Publish to Queue| I
  I --> J
  J -->|Send Notification| A
  J -->|Store in S3| H
Diagram

3. API design

  • POST /notifications: Create a new notification.
  • GET /notifications/{user_id}: Retrieve notifications for a user.
  • PUT /preferences/{user_id}: Update notification preferences for a user.
  • GET /preferences/{user_id}: Retrieve notification preferences for a user.

4. Data model & storage

Datastores:

  • SQL DB: For storing user preferences and metadata about notifications.
  • Object Storage (S3): For archiving sent notifications.
  • Redis Cache: For quick access to user preferences.

Key Tables:

  • notifications: Stores notification metadata (id, user_id, type, status, timestamp).
  • user_preferences: Stores user notification preferences (user_id, preference_data).

Partition/Sharding Key:

  • Use user_id as the partition key to distribute load evenly across the database.

5. Deep dive

To ensure timely delivery and reliability, the system uses a combination of WebSockets and a distributed pub/sub messaging system (Kafka). WebSockets enable real-time communication, ensuring low-latency delivery of notifications. Kafka provides a reliable message queue that supports large fan-out events and ensures message ordering.

sequenceDiagram
  participant U as User Device
  participant N as Notification Service
  participant P as Preferences Service
  participant Q as Kafka
  participant W as Worker
  participant S as S3 Storage

  U->>N: Request Notification
  N->>P: Check Preferences
  P-->>N: Return Preferences
  N->>Q: Publish Notification
  Q->>W: Consume Notification
  W->>U: Send Notification
  W->>S: Archive Notification
Diagram

6. Scale, bottlenecks & trade-offs

Replication & Sharding:

  • Use database replication for high availability.
  • Shard databases by user_id to distribute load.

Caching:

  • Use Redis to cache user preferences, reducing database load and improving response times.

Single Points of Failure:

  • Use multiple instances of services and load balancers to avoid single points of failure.
  • Implement failover strategies for Kafka and database clusters.

Trade-offs:

  • Consistency vs. Availability: Opt for eventual consistency in notification delivery to ensure high availability.
  • Push vs. Pull: Use push notifications for real-time delivery, but allow pull for retrieving missed notifications.
  • Sync vs. Async: Asynchronous processing of notifications ensures scalability and responsiveness.

By leveraging these strategies, the notification system can deliver messages reliably and in real-time, even under high load conditions.

TechnicalEasyAtlassian

17. What is the purpose of using version control systems in software development?

Model answer

Purpose of Version Control Systems in Software Development

  1. Collaboration and Coordination - Version control systems (VCS) enable multiple developers to work on the same project simultaneously without overwriting each other's changes. This is crucial for team collaboration, especially in large-scale projects.
  2. History and Traceability - VCS maintains a complete history of changes made to the codebase, allowing developers to track who made changes, what changes were made, and why. This is essential for auditing and understanding the evolution of the project.
  3. Branching and Merging - Developers can create branches to work on features or bug fixes independently. Once completed, these branches can be merged back into the main codebase. This supports parallel development and helps manage different versions of the software.
  4. Backup and Recovery - VCS acts as a backup system by storing snapshots of the project at various points in time. In case of data loss or corruption, developers can restore previous versions of the code.
  5. Conflict Resolution - When multiple developers make changes to the same part of the code, VCS provides tools to resolve conflicts, ensuring that the final code integrates all changes correctly.
  6. Continuous Integration and Deployment - VCS is integral to CI/CD pipelines, automating testing and deployment processes. This ensures that code changes are continuously integrated and deployed, reducing the time to market and improving software quality.
  7. Consistency and Reliability - By maintaining a single source of truth, VCS ensures that all team members are working with the most up-to-date and consistent version of the code, enhancing the reliability of the software development process.
  8. Facilitating Code Reviews - VCS supports code review processes by allowing developers to review changes, comment on them, and suggest improvements before they are merged into the main codebase.

In summary, version control systems are essential for managing code changes, facilitating collaboration, ensuring code quality, and maintaining a reliable and efficient software development process. They provide the necessary tools to handle the complexities of modern software engineering, making them indispensable in any development environment.

TechnicalMediumAtlassian

18. What is the role of Jira in software development?

Model answer

  1. Role of Jira in Software Development

Jira is a powerful tool widely used in software development for project management and issue tracking. It plays a crucial role in facilitating communication and collaboration among team members, which is essential for the successful delivery of software projects.

  1. Key Functions of Jira
  • Issue Tracking: Jira allows teams to create, update, and track issues or tasks throughout the development lifecycle. This helps in maintaining a clear overview of the project's progress and identifying any bottlenecks.
  • Agile Project Management: Jira supports agile methodologies such as Scrum and Kanban. It provides features like sprint planning, backlog grooming, and burndown charts, which help teams manage their workflows efficiently.
  • Customization and Integration: Jira is highly customizable, allowing teams to tailor workflows, issue types, and fields to match their specific processes. It also integrates with a wide range of development tools, enhancing its functionality and enabling seamless data flow across systems.
  • Reporting and Analytics: Jira offers robust reporting tools that provide insights into team performance, project progress, and potential risks. These reports help in making informed decisions and improving team productivity.
  • Collaboration and Communication: By centralizing information and providing a platform for discussion, Jira enhances team collaboration. Team members can comment on issues, attach files, and mention colleagues, ensuring everyone is aligned and informed.
  1. Benefits of Using Jira
  • Improved Transparency: Jira provides visibility into the status of tasks and projects, helping teams stay on track and stakeholders stay informed.
  • Enhanced Productivity: By automating repetitive tasks and streamlining workflows, Jira allows teams to focus on high-value activities, thereby improving productivity.
  • Better Risk Management: With its reporting and tracking capabilities, Jira helps teams identify and mitigate risks early in the development process.
  • Scalability: Jira can scale with the organization, supporting small teams as well as large enterprises with complex project management needs.

In summary, Jira is an essential tool in software development that supports effective project management, facilitates communication, and enhances team productivity through its comprehensive set of features tailored for agile methodologies.

TechnicalMediumAtlassian

19. Describe the importance of automated testing in software development.

Model answer

Importance of Automated Testing in Software Development

  1. Ensures Code Quality and Reliability - Automated testing helps maintain high code quality by catching bugs and issues early in the development cycle. This leads to more reliable software, as defects are identified and resolved before reaching production.
  2. Facilitates Continuous Integration and Continuous Deployment (CI/CD) - Automated tests are integral to CI/CD pipelines, allowing for rapid and reliable software deployment. They enable developers to integrate code changes frequently and with confidence, knowing that automated tests will verify the stability of the codebase.
  3. Improves Development Efficiency - By automating repetitive testing tasks, developers can focus on writing new features and improving existing ones. Automated tests run faster than manual tests, providing quick feedback and reducing the time spent on regression testing.
  4. Enhances Maintainability - Automated tests serve as documentation for the codebase, illustrating how different parts of the system should behave. This makes it easier for new developers to understand the system and for existing developers to make changes without introducing new bugs.
  5. Supports Agile Development Practices - Agile methodologies emphasize iterative development and frequent releases. Automated testing supports this by ensuring that each iteration maintains the integrity of the software, allowing for rapid adaptation to changing requirements.
  6. Reduces Human Error - Manual testing is prone to human error, especially when dealing with complex scenarios or large volumes of tests. Automated testing eliminates this risk by executing tests consistently and accurately every time.
  7. Facilitates Test Coverage and Scalability - Automated testing allows for extensive test coverage, including edge cases that might be overlooked in manual testing. As the software scales, automated tests can be expanded to cover new functionalities without a proportional increase in testing time.
  8. Enables Test-Driven Development (TDD) - Automated testing is a cornerstone of TDD, where tests are written before the code itself. This approach ensures that the code meets the specified requirements from the outset and encourages simple, clean designs by focusing on the necessary functionality.

In summary, automated testing is crucial for maintaining software quality, supporting agile practices, and ensuring efficient and reliable software development. It aligns with principles like KISS by simplifying the testing process and focusing on essential functionality, thereby enhancing maintainability and reducing complexity.

TechnicalMediumAtlassian

20. What is the main purpose of using a continuous integration system in software development?

Model answer

The main purpose of using a continuous integration (CI) system in software development is to improve the quality and efficiency of the development process by automating the integration of code changes from multiple contributors into a shared repository. Here are the key aspects of why CI is essential:

  1. Automated Testing: - CI systems automatically run tests on new code commits. This ensures that any changes do not break existing functionality, maintaining code quality and reliability.
  2. Early Bug Detection: - By integrating code frequently and running tests automatically, CI helps in identifying bugs early in the development cycle. This reduces the cost and effort required to fix issues compared to finding them later in the process.
  3. Improved Collaboration: - CI facilitates better collaboration among developers by integrating code changes frequently. This reduces integration problems and conflicts, enabling teams to work more cohesively.
  4. Faster Feedback: - Developers receive immediate feedback on their code changes, allowing them to address issues promptly. This accelerates the development process and improves the overall efficiency of the team.
  5. Reduced Integration Risk: - Frequent integration reduces the risk of integration problems that can occur when code changes are merged after long periods. This leads to smoother and more predictable releases.
  6. Consistent Build Process: - CI systems ensure a consistent build process by automating the compilation and testing of code. This reduces human error and ensures that the build process is repeatable and reliable.
  7. Enhanced Code Quality: - By enforcing coding standards and running static analysis tools, CI systems help maintain high code quality across the development team.
  8. Faster Time to Market: - With automated testing and integration, teams can release features and updates more frequently and with greater confidence, reducing the time to market for new products and features.

In summary, a continuous integration system is a critical component in modern software development that enhances collaboration, improves code quality, and accelerates the delivery of software products. By automating the integration and testing process, CI systems help teams maintain a high standard of software quality while enabling rapid development cycles.

Practice these out loud, don't memorise them

Reading an answer is not the same as being able to give one under pressure. ChannelPulse plays the interviewer, asks the follow-ups, and scores each answer with feedback and a model answer so you can hear the gap between what you said and what lands.

Get ChannelPulse Browse all questions