OfferTransform Your Career with Expert-Led IT Training. Flat discounts active!Explore Now
OnlineITGuru Logo
Software Development

MongoDB Indexes: What They Are and Why They Matter

Last updated on Sep 29, 2026

Copy Link:
MongoDB Indexes: What They Are and Why They Matter

The initial stage of an application is characterized by an insignificant data size. It is not difficult to find databases that perform exceedingly fast at that stage of development. Each of the collections within the application contains a few hundreds or thousands of documents. Single-digit millisecond query execution time and only a small fraction of memory usage ensure a hassle-free experience.

However, the perception of a fast system disappears over time. The datasets enlarge and exponentially increase in size, moving from thousands of records to hundreds of millions. The RAM capacity is growing insufficient, which means that the application needs to perform requests to the persistent storage on a continuous basis.

They are the basic elements that make it possible for the database to navigate, filter, and process the necessary data. In the absence of indexes, any query would be just a slow, unending process of processing and reading all the bytes on the hard disk. However, with indexes in place, every query results in locating documents instantaneously.

Understanding the concepts of MongoDB indexes, their workings, and how to create them effectively is crucial for developing highly scalable document-oriented applications. For those looking forward to learning the basics of these concepts and becoming an expert at them, taking structured mongodb classes or a complete mongodb course offered by OnlineITGuru is the way to go.

Search Without an Index

In order to understand the importance of indexes, one has to consider how below efficient searching would be in their absence.

Imagine that you are in a library with ten million books. There is no arrangement of books by titles, genres, authors, or other criteria. The only factor which is taken into account while putting new books is the availability of free space on the shelf. Would you be able to find a specific biography published in 1984?

The only option available to the librarian is to proceed to the first shelf on the first floor, take the first book from it, see its title and date of publication, return it and take the second book and repeat this process until the desired book is found or all the shelves have been checked.

In database terms, such a direct operation is called a collection scan or COLLSCAN.

Collection Scan's Mechanism and Price

If a client requires documents that meet certain conditions from a non-indexed collection, MongoDB doesn't know where those documents are stored. The storage layer must read all documents of the collection successively from the beginning to the end.

This operation requires a significant amount of processing:

  • Disk Input/Output Exhaustion: Collections can outlive their memory in production. Collection scans can force the operating system to keep calling from solid-state drives or spinning disk to memory. Accessing disk memory is much slower than performed by the CPU and thus becomes a bottleneck shortly after the start of the process.

  • Cache Eviction and Churn: MongoDB uses the default storage engine that is called WiredTiger, which uses an in-memory cache of frequently-used data pages. The collection scan pulls gigabytes of unneeded cold data from the cache and evicts warm, frequently-used documents from the cache. As a result, unrelated queries run more slowly since they had their cached working set dumped on the disk.

  • CPU Saturation: Since documents are converted into data, the CPU must check every single field in the documents. When there are millions of documents involved this could lead to an exhausting demand for processing power, leaving nothing for other ongoing functions to run in parallel.

  • System Concurrency Starvation: Because MongoDB applies the document level concurrency technique, the numerous reads performed as a result of unindexed queries make worker threads busy for a long period.

When scanning a collection we can be sure that time needed to finish the query will change directly in relation to the size of the collection. For instance, if the collection is twice as large, the query will take twice as long to execute. In this case, a few unindexed queries being executed at the same time may lead to the database becoming unusable in the conditions of heavy load.

What Are MongoDB Indexes?

In essence, an index is a supporting data structure that contains an organized and sorted subset of the information in a collection.

As an analogy to a library, an index can be considered to be the beautiful card archive or the index found at the back of a reference book. Instead of reading every page of a book to find information about a certain thing, you look at the index for that term, and find what pages to read.

Having created an index in MongoDB, the system selects and sorts the value of specified fields within each document of your collection and places the index entry into the absolutely correct location, i.e., another index field called record ID indicates where that record is stored in the cache or disk.

When you query records based on a particular field with the expected value or within a specific value range, MongoDB does not read the raw collection. Rather, it accesses a compact, already sorted index structure. After finding the index entries, MongoDB fetches only those documents whose addresses are stored in these indexes.

This dramatically reduces the time it takes for queries. For example, with a collection of ten million documents, an unindexed query will scan those ten million entries, while an indexed lookup would make no more than thirty comparisons.

The Engine Room: Performing Index Management with WiredTiger

In order to appreciate the role of indexing in the proper manner, it is necessary to learn how the engine that stores data assembles and processes this information.

MongoDB uses WiredTiger as its default engine—this means that collections and indexes are stored in separate dedicated files on disk. Although documents themselves remain compressed and stored as data pages, indexes are structured as self-balancing trees, mostly B-trees.

The Principle Behind B-tree

The B-tree is a tree structure that is organized hierarchically and created for storage systems that process large volumes of data. Unlike a traditional binary search tree that permits each node to have just two children, a node within a B-tree can have hundreds or thousands of children at a time. This feature, known as high-ordering branching, keeps the tree shallow. Consequently, the distance between the root node of the tree and the leaves is quite short because of only three or four levels separating it from millions of records.

There are several important properties that the B-tree maintains:

  • Sorted key arrangement: All keys kept in the nodes come in a completely sorted format, allowing the node to perform binary search functions.

  • Balanced level: All leaf nodes remain at the same level from the root. This allows all operations (lookup, insertion, and deletion) to have predictable logarithmic performance.

  • Range traversal: Since the entries are sorted in the leaves, the storage engine can start from any range (for example, the transactions that are after a specific date) and then simply move forward or backward without having to backtrack through the root

Compressing Prefixes and Cost of Memory

Memory being the key resource for every database system, WiredTiger uses prefix compression for its index keys both in memory and on the disk. When many adjacent entries in the index share either beginning characters or prefixes, the system saves just the prefix and records only unique suffixes for the rest of the entries.

The high level of compression permits creating very big indexes but still not make the index bigger than the WiredTiger RAM cache. Thus, if the required data fit in memory, the queries can be processed at the speed of transferring data via electrical bus instead of reading it from the disk.

Diversity of Index Types

The document model of the MongoDB is characterized by broad flexibility with structure allowing to create nested subdocuments and various records and arrays. In line with such a flexible schema model, the MongoDB has many different index types created for different types of structures and queries.

Single Field Index

The single field index serves as the essential component of the indexing within databases wherein a single attribute of every document is taken into account and placed into a sorted B-tree.

Every MongoDB collection possesses an immutable single field index that operates in relation to the usual primary key field of any collection, which means that primary key values are searched for instantly. Additionally, it blocks any client application from inserting documents with duplicate keys.

It is noteworthy that the developers may create other single field indexes based on any property of the document (for instance, an email address, username, or date of creation). When it comes to single field indexing, sort order, which can be ascending or descending, does not really affect the speed of queries because it is possible to access single key B-tree both ways.

Compound Indexes

Today, modern applications seldom filter on a single attribute. A case in point would be an e-commerce application that searches for all the orders by a specific customer along with their fulfillment status, and orders the results based on the order date.

A compound index consists of several independent fields combined into one single index. The order in which the fields are arranged plays a crucial role; actually, the index is organized based on the values in the first field; when field one has several records with the same value, the values are sorted in accordance with the second field, etc.

As a result of this strict hierarchy, a compound index supports the following types of queries:

  • The exact combination of all indexed fields.

  • Any possible prefix of these fields.

For example, an index built on 3 fields (e.g. country, state and city) will work well for queries filtering on country, or queries filtering on country and state, or queries filtering on country, state and city. However, it cannot efficiently support a query that filters on city or state alone, because the underlying B-Tree is partitioned by country at its highest level.

Also, the sort direction of a compound index is important, unlike single field indexes. If the query asks for data sorted by one field ascending and another descending, then the compound index must be created with matching or exact inverse directional polarity to satisfy that sort without having to do an expensive in-memory operation.

Several Key Indexes

Native arrays MongoDB supports native arrays, one of its defining features. A document for a published article might have an array of topical tags, or a user profile might have a list of associated contact phone numbers.

MongoDB automatically creates a multikey index if you create an index on a field that contains an array. Instead of creating a single index entry for the whole array, the engine unpacks the array and creates separate, distinct index entries for each of the items in the array.

For example, a document with a tags array of five values will create five distinct pointers in the B-Tree index. This enables searches that look for documents containing a specific tag to find those documents instantaneously.

But this flexibility is bounded by the following restriction: to prevent combinatorial explosion of index entries, a compound multikey index cannot index more than one array field in the same document. If two parallel array fields were to be indexed together, the engine would have to build the Cartesian product of both arrays, consuming huge amounts of storage and memory.

Indexes

Standard B-Tree indexes support exact value comparisons and prefix matches. They are not effective at searching for arbitrary words scattered in free-form text, such as product descriptions, blog posts or customer reviews.

  • Text index uses natural language processing techniques to analyse the content of a string. When creating a text index, MongoDB:

  • Removes punctuation

  • Splits text into individual lexical tokens or words.

  • Removes language specific stop words (e.g. common articles and prepositions such as "the", "an", or "with").

Uses language-aware stemming algorithms to reduce words to their linguistic roots (e.g. reducing "running", "ran" and "runs" to the common base word "run").

This index enables users to perform keyword searches over large collections of text without the need for full external search engine infrastructure for simpler search workloads.

Spatial Indexes

Applications that are location-aware need to be able to do spatial calculations such as finding the nearest coffee shops, verifying whether a delivery vehicle is within a specific service radius, or customers within a regional boundary polygon.

MongoDB has geospatial indexes that are specialised:

  • Planar (2D) Indexes: For flat, Euclidean geometry, such as indoor floor plans or game grids.

  • Spherical indexes (2dsphere): Built for modelling the round, three-dimensional geometry of the planet Earth.

Spherical indexes map coordinates to a virtual globe taking into account the curvature of the Earth for distance, intersection and containment calculations. These enable the application to perform complex spatial queries with very high computational efficiency.

Hashed Indexes

In a large distributed environment with sharded clusters, data should be partitioned evenly across physical database instances. If documents are shard by a monotonically increasing key (e.g. an auto-incrementing timestamp), all new write operations will hit a single shard and create a serious bottleneck.

A hashed index takes a hash ( based on MD5 ) of the value of the specified field, and indexes that hash. What this does is take sequential or near sequential inputs and turn them into a nice, evenly distributed pseudo random series of numbers. A hashed index, used as a shard key, ensures an even distribution of both writes and data storage across all of the cluster’s shards .

Indexes Wildcard

On highly polymorphic schemas (for example, arbitrary user-defined metadata dictionaries, product catalogues where every category has twenty completely unique attributes, or unstructured sensor payloads) it is not possible to pre-define indexes on every possible field.

Wildcard indexes provide a solution to this problem, allowing developers to define an index pattern that encompasses all arbitrary sub-fields within a given subdocument, or within the entire document. The engine dynamically indexes any scalar field it finds in the targeted path allowing for flexible index coverage without having to rigidly declare indexes upfront for hundreds of different fields.

Architectural Modifiers and Behavioural Properties

Beyond the structural classification, MongoDB indexes can be configured with behavioural flags that alter the way of storing, validating and pruning entries.

Unique index

Sure , indexes are most often lauded for optimising reads , but they also act as enforceable data integrity guards . A unique index tells the database to reject any insert or update that would make duplicate values for the indexed key.

This moves data integrity validation from the application layer into the database engine, removing race conditions where multiple application threads may check for existence and attempt to insert the same records simultaneously.

Be careful about missing fields when using unique constraints in a document database. In MongoDB, if a document is missing an indexed property, the database considers that missing property to have a value of null. A second document without that property will be therefore rejected as a duplicate null entry unless it is combined with sparse or partial modifiers.

Partial indexes

In many applications, only a subset of a collection is actively queried. Imagine an order tracking system with tens of millions of historical orders over the last decade. Most of these orders are marked as completed or archived and rarely searched. But in contrast, there are only a few thousand active/pending/out-for-delivery orders at a time, but those active orders get almost all the read traffic of the application.

You waste precious memory and disc space building an index of all tens of millions of records. A partial index would be a solution here, with a declarative filter rule. It indexes only documents that pass some criteria.

Here you can tell the partial index to include only documents where the order status equals active. The index generated is small, fits easily into system RAM, costs virtually nothing to update when historic documents are left alone, and performs lightning fast searches for all high frequency operational queries.

Sparse Indexes

Sparse indexes are a simpler type of conditional indexing that came earlier. Sparse index means that only the documents which do have the indexed property are included in the B-Tree. Documents where the field does not exist at all are completely excluded.

Partial indexes have largely replaced sparse indexes, although the latter are useful for dealing with schema variations. Partial indexes are a strict superset of sparse index capabilities, allowing developers to define expressive, multi-faceted criteria, instead of just checking for field existence.

Indexes with TTL

Modern systems produce massive amounts of ephemeral data: user authentication sessions, ephemeral verification tokens, API rate-limiting records, diagnostic trace logs. If left unchecked these collections grow without end consuming storage and degrading performance.

A Time-to-Live (TTL) index is a special single field index on a date-type field that automatically deletes documents older than a given age threshold.

Under the hood, MongoDB runs an asynchronous background thread that checks TTL indexes periodically (typically once a minute) and deletes expired documents automatically. This removes the need for external cron jobs or custom garbage-collection worker services and moves data lifecycle governance directly to the database layer.

Concealed Indexes

Changing indexes in high throughput production environments is always a risk. Removing an index that later becomes critical can immediately put an application in collection-scan paralysis.

A hidden index is an index fully maintained by the storage engine on inserts and updates, but which the query planner can deliberately not see.

This allows engineers to safely test the effect of not having an index: administrators can hide the index and see if it causes queries to degrade or fall back to inferior plans. In case of any problem, the index can be unhidden immediately without the high cost of rebuilding the entire tree structure from scratch.

Beyond Indexing

The knowledge of the inner workings of indexes like balancing in B-Tree leafs and evicting items in cache by WiredTiger is just a small fraction of designing today’s modern document databases. In any real-world architecture of an enterprise organization, there are many other concepts to consider in addition to what we’ve discussed so far, such as replica sets, write concerns, sharding solutions, optimization in aggregation pipelines, and transactions.

If you are interested in taking your database engineering skills from theory to practice, I suggest considering a full mongodb learning path. The mongodb full course at OnlineITGuru will provide you with practical lab experience, real-world examples, and hands-on projects that will help you develop skills in CRUD patterns and more.

Why Choose Us

Master Your Future with OnlineITGuru

We don't just provide courses; we build careers. From expert-led live training to dedicated placement support, discover why thousands of professionals trust us for their digital transformation journey.

200+

Partner Companies

$120K

Highest Package

75%

Average Hike

98%

Placement Rate

Reliable Career Partners

Google
Microsoft
Amazon
Meta
Netflix
Apple