A geospatial foundation model (GFM), also referred to as an Earth observation foundation model, represents a paradigm shift in Earth system modeling by utilizi
Type of artificial intelligence model trained on Earth observation and geoscientific data
A geospatial foundation model (GFM), also referred to as an Earth observation foundation model, represents a paradigm shift in Earth system modeling by utilizing data-centric artificial intelligence (DCAI) to train on petabytes of structured and unstructured geoscientific data.[1] Unlike conventional, task-specific deep learning systems, GFMs use massive cross-disciplinary datasets to model the Earth's interactive and complex dynamics, offering highly flexible task specifications, multimodal inputs/outputs, and advanced geoscientific knowledge representation.[citation needed]
Despite these early successes, traditional artificial intelligence approaches face prominent challenges:
Data Dependency: Traditional models heavily rely on extensive, high-quality annotated datasets, which are often limited or incomplete due to the variable nature of Earth systems.[14]
Limited Generalization: The capacity of static AI models to generalize often degrades significantly when encountering novel or distinct geological environments outside their training context.[15][16]
Lack of Interpretability: Opaque model outputs challenge the transparency required for geoscientists to firmly establish trust in AI insights.[17][18]
To address these limitations, foundation models (FMs) have transitioned into the geosciences from computer vision and natural language processing.[19] FMs offer emergent properties such as zero-shot adaptability and contextual reasoning derived from massive neural network scale and self-supervised learning on vast datasets.[20][21] Early efforts to pioneer GFMs have spanned multiple specializations[22][23][24] though broad adoption remains constrained by multi source data complexity and the inherent constraints of contemporary AI architectures.[22][23][24][25]
Architecture and core mechanisms
GFM architectures primarily leverage the modeling capacity of the Transformer architecture, using self-attention mechanisms to process long-range dependencies and contextual relationships,[26] mitigating the historical constraints of recurrent neural networks (RNNs) and convolutional neural networks (CNNs).[27][28]
Transformer full architecture
Core transformer frameworks
Vanilla Transformers: Composed of encoder and decoder blocks featuring multihead self-attention (MHSA) and position-wise feed-forward layers, to capture contextual sequence vectors uniformly across text length.[29][30]
Vision Transformers (ViTs): Adapt the architecture to computer vision tasks by dividing input visual data into localized 16 by 16 patches, flattening them with linear projections, and combining them with positional encodings through an encoder framework.[30] Notable standard variations include TNT,[31] and PVT.[32]
Vision-Language Transformers: Fuse multimodal visual and textual streams via unified attention mechanisms to capture cross-modal associations for visual question answering (VQA) and image captioning, historically building on models like VisualBERT,[33] Uniter,[34] and OSCAR[35] which combine BERT text processing with Faster-RCNN object detection pipelines.[33][34][35]
Pretraining taxonomy
GFMs acquire general-purpose weights and broad structural representations through expansive training methods:[36]
Supervised Pretraining: Utilizes large-scale labeled datasets to establish model pillars, historically scaled up through CNN backbones like AlexNet,[37] VGG,[38] and ResNet[39] over ImageNet or equivalent repositories[40] and through Transformer platforms like BERT,[41] RoBERTa,[42] and ALBERT.[43]
Self-Supervised Pretraining (SSL): Acts as the cornerstone for modern large-scale systems by using pretext tasks to derive implicit pseudo-labels from unlabeled inputs, minimizing human annotation demands.[44][45][46] SSL paradigms include:
Generative Pretraining: Captures underlying statistical parameters by reconstructing obscured or generated data fields, utilizing frameworks like Variational Autoencoders (VAEs),[47] Context Encoders for inpainting,[48] Masked Autoencoders (MAEs),[49] and Generative Adversarial Networks (GANs) such as BigGAN[50] and SRGAN.[51]
Contrastive Pretraining: Optimizes feature representations by bringing matching or augmented data pairs closer in embedding space while repelling dissimilar samples. Core contrastive workflows are anchored by negative sampling (e.g., SimCLR,[52] MoCo iterations[53][54]), data clustering (e.g., DeepCluster,[55] SwAV[56]), teacher-student knowledge distillation (e.g., BYOL,[57] DINO variants[58]), and cross-correlation matrix redundancy reduction (e.g., Barlow Twins).[59]
Predictive Pretraining: Relies on concrete pretext calculations such as relative spatial patch alignment,[60] geometric spatial transformations (e.g., RotNet[61]), or spectral color mapping (e.g., Image Colorization tasks[62]).
Hybrid Pretraining: Integrates complementary strategies—such as cross-modal contrastive training matched with autoregressive decoding or self-supervised generation paired with downstream oversight—exemplified by multimodal models like CLIP[63] and DALL-E.[64]
Adaptation strategies
Pretrained foundational backbones are tuned for specific target geoscientific tasks or localized tracking using parameter-efficient frameworks:[65][66]
Fine-Tuning: Alters the complete weight configuration using task-specific datasets, often employing smaller learning rates to preserve pretrained global representations.
Prompt Tuning: Adapts models to downstream contexts without altering underlying weights. In visual-spatial settings, these are deployed as vision-driven prompts (e.g., VPT,[67] DePT,[68] PVIT[69]), language-driven text prompts (e.g., CoOp,[70] PLOT[71]), or dual vision-language prompts (e.g., UPT,[72] DPT,[73] TPT[74]).
Adapter Tuning: Adds compact, trainable parameters to a frozen base network. These can follow a sequential setup within forward architectures (e.g., Res-adapt,[75] DAN,[76] LST,[77] Conv-Adapter[78]), run parallel to traditional sublayers (e.g., ViT-Adapter,[79] PESF-KD,[80] AdaptMLP[81]), or blend architectures dynamically inside multihead attention systems as mix adapters (e.g., PATT,[82] ETT,[83] PALT,[84] VQT[85]).
Parameter Tuning: Modifies parameter groups via reparameterization paths, isolating localized weight adjustments (e.g., LoRA,[86] DyLoRA[87]), targeting structural bias updates (e.g., Bitfit,[88] AdapterBias[89]), or modulating weight and bias parts uniformly (e.g., SSF[90]).[86][87][90]
Notable models
Advancements in geoscientific foundation models map across distinct modality systems, expanding past foundational language architectures like GPT variants,[91] BERT versions[42] and general text reasoning techniques.[92]
Large Language Models (LLMs) for geoscience
GeoBERT: Developed by retraining standard BERT models over 20 million internal geological and geoscientific records to perform field-specific query answering and localized textual summarization.[93]
BERT-E: Built through targeted domain transfer learning from SciBERT to specialize in earth science keyword classification.[94]
K2: Pioneered as a dedicated geoscience-specific language foundation model utilizing the unique instruction-driven GeoSignal training dataset and the GeoBench verification collection.[95]
CnGeoPLM & GeoBERTSegmenter: Optimized for non-English records; CnGeoPLM handles geological entity and relationship extraction over a specialized Chinese text corpus (GeoCorpus),[96] while GeoBERTSegmenter resolves complex geological term word tokenization.[97]
OceanGPT: Formulated as a dedicated language engine for oceanography and marine sciences using the DoInstruct data method and measured via the OceanBench platform.[98]
GeoGalactica: A massive specialized 30-billion parameter geoscience model trained explicitly on the extensive GeoCorpus text resource.[99]
Geospatial Copilots and GIS Engines: Implementations like GPT4GEO, GEOGPT, BB-GeoGPT, and GeoLLM utilize text, geometry, and location engines (such as OpenStreetMap) to evaluate geospatial questions, measure socioeconomic livelihoods, and build realistic automation assistants.[19][100][101][102][103][104][105][106]
Large Vision Models (LVMs) for geoscience
RingMo Series: Introduces a scalable remote sensing model framework driven by masked image modeling over 2 million spatial scenes.[11] Variants include plain ViT models adjusting rotated window dimensions,[12] spatiotemporal sequence trackers (RingMo-Sense),[107] and resource-optimized edge networks (RingMo-Lite).[108]
Segment Anything Model (SAM) Remote Sensing Adaptations: Adaptations and automated tracking pipelines mapping SAM to spatial datasets include RingMo-SAM,[109] large-scale segment annotation generation tools like SAMRS[110] and RSPrompter,[111] boundary-constrained frameworks,[112] and targeted asset identifiers like GeoSAM for mobility networks.[113]
Prithvi: An open-source spatial transformer backbone engineered jointly by NASA and IBM, pretrained using more than 1 terabyte of multi-spectral Harmonized Landsat Sentinel-2 (HLS) satellite observations.[114]
Atmospheric Vision Models: Foundational frameworks applied directly to weather systems and climate dynamics include Climax,[115] W-MAE,[116] FourCastNet,[117] GraphCast,[118] Fengwu,[119] and Pangu-Weather.[120]
SkySense: A multi-modal visual architecture showing state-of-the-art generalization across global earth image understanding benchmarks.[118][119][120][121]
Large Vision-Language Models (LVLMs) for geoscience
GeoChat: Introduced as the first grounded remote sensing visual-language model delivering multi-task conversational features over coordinate inputs.[122]
EarthGPT & SkyeyeGPT: Multi-modal frameworks deployed for open-set visual scene analysis, multi-sensor imagery captioning, and conversational instruction response.[123][124]
RemoteCLIP & CLIP-RS: Adapt cross-modal text-image contrastive training paradigms for remote sensing, using unmanned aerial vehicle views and massive image-caption datasets.[125][126]
Lightweight and Multimodal RSVQA Systems: Architectures like LIT-4-RSVQA[127] and prompt-guided visual platforms[128] track multi-spectral queries while evaluating structural dataset biases.[129][130]
Foundational model agents
Building on autonomous multi-module agent definitions (comprising perception pipelines, brain modules for storage and decision making, and tool-action vectors),[131][132] specialized geospatial software agents have emerged:
RS-ChatGPT: Implements a prompt-driven tool orchestration wrapper enabling ChatGPT to map, segment, and count remote sensing data fields by deploying external vision library scripts.[133]
Change-Agent: Fuses multi-level change tracking models with text logic to detect, count, and explain terrain alterations over temporal satellite spans.[134]
RS-Agent: Merges expert knowledge libraries with text and image processing systems to resolve open-ended operational queries in spatial engineering tasks.[135]
Applications
GFMs are deployed across diverse observation vectors spanning space, air, terrestrial ground systems, and deep oceans:
Remote sensing, environmental monitoring, and land use
The application of spatial vision platforms significantly enhances pixel classification over hyper-spectral and multi-spectral arrays, automating geographic land-use and land-cover (LULC) segmentation workflows without labor-intensive feature manual curation.[136][24] These architectures track long-term deforestation trends, model carbon biomass reserves, isolate vegetation changes, and manage building extraction footprints or structural urban modifications over time.[137][138]
Climate dynamics and weather forecasting
GFMs function as unified multi-source processing centers, assimilating disjointed observation data from weather stations, high-resolution radars, and orbiting satellites to generate contextually unified semantic representations.[139] Deployed predictive networks downscale coarse global grid targets, measure regional drought extensions, track severe typhoon pathways, and enhance the tracking velocity of extreme precipitation events.[140][141][142]
Terrestrial geophysics and seismology
In terrestrial ground settings, deep spatial architectures manage high-dimensional waveforms, historical subsurface geological profiles, and live sensor measurements. This assists in modeling intricate crustal deformations, updating geological hazard maps, mapping deep mineral or petrochemical reserves, and evaluating earthquake behaviors, seismic waveforms, or chaotic landslide hazards.[143][144][145][146][147]
Oceanography and marine exploration
GFMs support interactive three-dimensional visualization networks mapping sub-surface, biological, and environmental observation streams. Deployed across intermediate and deep midocean zones (from 100 to 1,000 meters deep), these systems leverage intelligent sensor fusion to explore sparse marine coordinates, evaluate temperature fields, simulate fluid dynamics, and optimize marine conservation choices within Deep Blue AI platforms.[148][149]
References
^Zhang, Hao; Xu, Jin-Jian; Cui, Hong-Wei; Li, Lin; Yang, Yaowen; Tang, Chao-Sheng; Boers, Niklas (December 2025). "When Geoscience Meets Foundation Models: Toward a general geoscience artificial intelligence system". IEEE Geoscience and Remote Sensing Magazine. 13 (4): 79–118. arXiv:2309.06799. Bibcode:2025IGRSM..13d..79Z. doi:10.1109/MGRS.2024.3496478.
^Anantrasirichai, N.; Biggs, J.; Albino, F.; Bull, D. (September 2019). "A deep learning approach to detecting volcano deformation from satellite imagery using synthetic datasets". Remote Sensing of Environment. 230 111179. arXiv:1905.07286. Bibcode:2019RSEnv.23011179A. doi:10.1016/j.rse.2019.04.032.
^Hess, Philipp; Drüke, Markus; Petri, Stefan; Strnad, Felix M.; Boers, Niklas (3 October 2022). "Physically constrained generative adversarial networks for improving precipitation fields from Earth system models". Nature Machine Intelligence. 4 (10): 828–839. Bibcode:2022NatMI...4..828H. doi:10.1038/s42256-022-00540-1.
^Mitsui, Takahito; Boers, Niklas (July 2021). "Seasonal prediction of Indian summer monsoon onset with echo state networks". Environmental Research Letters. 16 (7): 074024. Bibcode:2021ERL....16g4024M. doi:10.1088/1748-9326/ac0acb.
^Weyn, Jonathan A.; Durran, Dale R.; Caruana, Rich (August 2019). "Can Machines Learn to Predict Weather? Using Deep Learning to Predict Gridded 500-hPa Geopotential Height From Historical Weather Data". Journal of Advances in Modeling Earth Systems. 11 (8): 2680–2693. Bibcode:2019JAMES..11.2680W. doi:10.1029/2019MS001705.
^ abSun, Xian; Wang, Peijin; Lu, Wanxuan; Zhu, Zicong; Lu, Xiaonan; He, Qibin; Li, Junxi; Rong, Xuee; Yang, Zhujun; Chang, Hao; He, Qinglin; Yang, Guang; Wang, Ruiping; Lu, Jiwen; Fu, Kun (2023). "RingMo: A Remote Sensing Foundation Model With Masked Image Modeling". IEEE Transactions on Geoscience and Remote Sensing. 61: 1–22. Bibcode:2023ITGRS..6194732S. doi:10.1109/TGRS.2022.3194732.
^ abWang, Di; Zhang, Qiming; Xu, Yufei; Zhang, Jing; Du, Bo; Tao, Dacheng; Zhang, Liangpei (2023). "Advancing Plain Vision Transformer Toward Remote Sensing Foundation Model". IEEE Transactions on Geoscience and Remote Sensing. 61: 1–15. Bibcode:2023ITGRS..6122818W. doi:10.1109/TGRS.2022.3222818.
^Ding, Lei; Zhu, Kun; Peng, Daifeng; Tang, Hao; Yang, Kuiwu; Bruzzone, Lorenzo (2024). "Adapting Segment Anything Model for Change Detection in VHR Remote Sensing Images". IEEE Transactions on Geoscience and Remote Sensing. 62: 1–11. arXiv:2309.01429. Bibcode:2024ITGRS..6268168D. doi:10.1109/TGRS.2024.3368168.
^Karpatne, Anuj; Ebert-Uphoff, Imme; Ravela, Sai; Babaie, Hassan Ali; Kumar, Vipin (2019). "Machine Learning for the Geosciences: Challenges and Opportunities". IEEE Transactions on Knowledge and Data Engineering. 31 (8): 1544–1554. arXiv:1711.04708. Bibcode:2019ITKDE..31.1544K. doi:10.1109/TKDE.2018.2861006.
^Aslam, Rana Waqar; Shu, Hong; Javid, Kanwal; Pervaiz, Shazia; Mustafa, Farhan; Raza, Danish; Ahmed, Bilal; Quddoos, Abdul; Al-Ahmadi, Saad; Hatamleh, Wesam Atef (February 2024). "Wetland identification through remote sensing: Insights into wetness, greenness, turbidity, temperature, and changing landscapes". Big Data Research. 35 100416. Bibcode:2024BDR....3500416A. doi:10.1016/j.bdr.2023.100416.
^Chen, Keyan; Chen, Bowen; Liu, Chenyang; Li, Wenyuan; Zou, Zhengxia; Shi, Zhenwei (2024). "RSMamba: Remote Sensing Image Classification With State Space Model". IEEE Geoscience and Remote Sensing Letters. 21: 1–5. arXiv:2403.19654. Bibcode:2024IGRSL..2107111C. doi:10.1109/LGRS.2024.3407111.
^Toms, Benjamin A.; Barnes, Elizabeth A.; Ebert-Uphoff, Imme (September 2020). "Physically Interpretable Neural Networks for the Geosciences: Applications to Earth System Variability". Journal of Advances in Modeling Earth Systems. 12 (9) e2019MS002002. arXiv:1912.01752. Bibcode:2020JAMES..1202002T. doi:10.1029/2019MS002002.
^ abBommasani, Rishi; Hudson, Drew A.; Adeli, Ehsan; Altman, Russ; Arora, Simran; Arx, Sydney von; Bernstein, Michael S.; Bohg, Jeannette; Bosselut, Antoine (2021). "On the Opportunities and Risks of Foundation Models". arXiv:2108.07258 [cs.LG].
^Värtinen, Susanna; Hämäläinen, Perttu; Guckelsberger, Christian (March 2024). "Generating Role-Playing Game Quests With GPT Language Models". IEEE Transactions on Games. 16 (1): 127–139. Bibcode:2024ITGam..16..127V. doi:10.1109/TG.2022.3228480.
^Hu, Yushi; Hua, Hang; Yang, Zhengyuan; Shi, Weijia; Smith, Noah A.; Luo, Jiebo (2023). "PromptCap: Prompt-Guided Image Captioning for VQA with GPT-3". 2023 IEEE/CVF International Conference on Computer Vision (ICCV). pp. 2951–2963. Bibcode:2023iccv.conf..277H. doi:10.1109/ICCV51070.2023.00277. ISBN979-8-3503-0718-4.
^ abLi, Xiang; Wen, Congcong; Hu, Yuan; Yuan, Zhenghang; Zhu, Xiao Xiang (June 2024). "Vision-Language Models in Remote Sensing: Current progress and future trends". IEEE Geoscience and Remote Sensing Magazine. 12 (2): 32–66. arXiv:2305.05726. Bibcode:2024IGRSM..12b..32L. doi:10.1109/MGRS.2024.3383473.
^ abFuller, Anthony; Millard, Koreen; Green, James R. (2022). "SatViT: Pretraining Transformers for Earth Observation". IEEE Geoscience and Remote Sensing Letters. 19: 1–5. Bibcode:2022IGRSL..1901489F. doi:10.1109/LGRS.2022.3201489.
^Wang, Wenhai; Xie, Enze; Li, Xiang; Fan, Deng-Ping; Song, Kaitao; Liang, Ding; Lu, Tong; Luo, Ping; Shao, Ling (2021). "Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions". 2021 IEEE/CVF International Conference on Computer Vision (ICCV). pp. 548–558. Bibcode:2021iccv.conf...61W. doi:10.1109/ICCV48922.2021.00061. ISBN978-1-6654-2812-5.
^ abLi, Liunian Harold; Yatskar, Mark; Yin, Da; Hsieh, Cho-Jui; Chang, Kai-Wei (2020). "What Does BERT with Vision Look At?". Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. pp. 5265–5275. doi:10.18653/v1/2020.acl-main.469.
^Liu, Xiao; Zhang, Fanjin; Hou, Zhenyu; Mian, Li; Wang, Zhaoyu; Zhang, Jing; Tang, Jie (2021). "Self-supervised Learning: Generative or Contrastive". IEEE Transactions on Knowledge and Data Engineering: 1. arXiv:2006.08218. doi:10.1109/TKDE.2021.3090866.
^Krizhevsky, Alex; Sutskever, Ilya; Hinton, Geoffrey E. (24 May 2017). "ImageNet classification with deep convolutional neural networks". Communications of the ACM. 60 (6): 84–90. doi:10.1145/3065386.
^Simonyan, Karen; Zisserman, Andrew (2014). "Very Deep Convolutional Networks for Large-Scale Image Recognition". arXiv:1409.1556 [cs.CV].
^Devlin, Jacob; Chang, Ming-Wei; Lee, Kenton; Toutanova, Kristina (2018). "BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding". arXiv:1810.04805 [cs.CL].
^ abLiu, Yinhan; Ott, Myle; Goyal, Naman; Du, Jingfei; Joshi, Mandar; Chen, Danqi; Levy, Omer; Lewis, Mike; Zettlemoyer, Luke (2019). "RoBERTa: A Robustly Optimized BERT Pretraining Approach". arXiv:1907.11692 [cs.CL].
^Lan, Zhenzhong; Chen, Mingda; Goodman, Sebastian; Gimpel, Kevin; Sharma, Piyush; Soricut, Radu (2019). "ALBERT: A Lite BERT for Self-supervised Learning of Language Representations". arXiv:1909.11942 [cs.CL].
^Jing, Longlong; Tian, Yingli (November 2021). "Self-Supervised Visual Feature Learning With Deep Neural Networks: A Survey". IEEE Transactions on Pattern Analysis and Machine Intelligence. 43 (11): 4037–4058. Bibcode:2021ITPAM..43.4037J. doi:10.1109/TPAMI.2020.2992393. PMID32386141.
^Chen, Xingyu; Zhao, Liye; Xu, Jiawen; Liu, Zhikang; Zhang, Chengxin; Xu, Luxiang; Guo, Ning (2024). "A Self-Supervised Contrastive Denoising Autoencoder-Based Noise Suppression Method for Micro Thrust Measurement Signals Processing". IEEE Transactions on Instrumentation and Measurement. 73: 1–17. Bibcode:2024ITIM...7341134C. doi:10.1109/TIM.2023.3341134.
^Caron, Mathilde; Bojanowski, Piotr; Joulin, Armand; Douze, Matthijs (2018). "Deep Clustering for Unsupervised Learning of Visual Features". arXiv:1807.05520 [cs.CV].
^Caron, Mathilde; Misra, Ishan; Mairal, Julien; Goyal, Priya; Bojanowski, Piotr; Joulin, Armand (2020). "Unsupervised Learning of Visual Features by Contrasting Cluster Assignments". arXiv:2006.09882 [cs.CV].
^Grill, Jean-Bastien; Strub, Florian; Altché, Florent; Tallec, Corentin; Richemond, Pierre H.; Buchatskaya, Elena; Doersch, Carl; Pires, Bernardo Avila; Guo, Zhaohan Daniel (2020). "Bootstrap your own latent: A new approach to self-supervised Learning". arXiv:2006.07733 [cs.LG].
^Oquab, Maxime; Darcet, Timothée; Moutakanni, Théo; Vo, Huy; Szafraniec, Marc; Khalidov, Vasil; Fernandez, Pierre; Haziza, Daniel; Massa, Francisco (2023). "DINOv2: Learning Robust Visual Features without Supervision". arXiv:2304.07193 [cs.CV].
^Radford, Alec; Kim, Jong Wook; Hallacy, Chris; Ramesh, Aditya; Goh, Gabriel; Agarwal, Sandhini; Sastry, Girish; Askell, Amanda; Mishkin, Pamela (2021). "Learning Transferable Visual Models From Natural Language Supervision". arXiv:2103.00020 [cs.CV].
^Ramesh, Aditya; Dhariwal, Prafulla; Nichol, Alex; Chu, Casey; Chen, Mark (2022). "Hierarchical Text-Conditional Image Generation with CLIP Latents". arXiv:2204.06125 [cs.CV].
^Lialin, Vladislav; Deshpande, Vijeta; Yao, Xiaowei; Rumshisky, Anna (2023). "Scaling Down to Scale Up: A Guide to Parameter-Efficient Fine-Tuning". arXiv:2303.15647 [cs.CL].
^Tu, Cheng-Hao; Mai, Zheda; Chao, Wei-Lun (2023). "Visual Query Tuning: Towards Effective Usage of Intermediate Representations for Parameter and Memory Efficient Transfer Learning". 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 7725–7735. Bibcode:2023cvpr.conf..746T. doi:10.1109/CVPR52729.2023.00746. ISBN979-8-3503-0129-8.
^ abHu, Edward J.; Shen, Yelong; Wallis, Phillip; Allen-Zhu, Zeyuan; Li, Yuanzhi; Wang, Shean; Wang, Lu; Chen, Weizhu (2021). "LoRA: Low-Rank Adaptation of Large Language Models". arXiv:2106.09685 [cs.CL].
^ abValipour, Mojtaba; Rezagholizadeh, Mehdi; Kobyzev, Ivan; Ghodsi, Ali (2022). "DyLoRA: Parameter Efficient Tuning of Pre-trained Models using Dynamic Search-Free Low-Rank Adaptation". arXiv:2210.07558 [cs.CL].
^ abLian, Dongze; Zhou, Daquan; Feng, Jiashi; Wang, Xinchao (2022). "Scaling & Shifting Your Features: A New Baseline for Efficient Model Tuning". arXiv:2210.08823 [cs.CV].
^Raffel, Colin; Shazeer, Noam; Roberts, Adam; Lee, Katherine; Narang, Sharan; Matena, Michael; Zhou, Yanqi; Li, Wei; Liu, Peter J. (2019). "Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer". arXiv:1910.10683 [cs.LG].
^Denli, Huseyin; Chughtai, Hassan A.; Hughes, Brian; Gistri, Robert; Xu, Peng (2021). "Geoscience Language Processing for Exploration". Abu Dhabi International Petroleum Exhibition & Conference D031S102R003. OnePetro. Bibcode:2021adip.conf07766D. doi:10.2118/207766-MS.
^Ramachandran, R.; Ramasubramanian, M.; Koirala, P.; Gurung, I.; Maskey, M. (2022). "Language Model for Earth Science: Exploring Potential Downstream Applications as well as Current Challenges". IGARSS 2022 - 2022 IEEE International Geoscience and Remote Sensing Symposium. pp. 4015–4018. Bibcode:2022igar.conf..989R. doi:10.1109/IGARSS46834.2022.9883682. ISBN978-1-6654-2792-0.
^Deng, Cheng; Zhang, Tianhang; He, Zhongmou; Xu, Yi; Chen, Qiyuan; Shi, Yuanyuan; Fu, Luoyi; Zhang, Weinan; Wang, Xinbing (2023). "K2: A Foundation Language Model for Geoscience Knowledge Understanding and Utilization". arXiv:2306.05064 [cs.CL].
^Ma, Kai; Zheng, Shuai; Tian, Miao; Qiu, Qinjun; Tan, Yongjian; Hu, Xinxin; Li, HaiYan; Xie, Zhong (December 2023). "CnGeoPLM: Contextual knowledge selection and embedding with pretrained language representation model for the geoscience domain". Earth Science Informatics. 16 (4): 3629–3646. Bibcode:2023EScIn..16.3629M. doi:10.1007/s12145-023-01112-6.
^Wei, Dongqi; Liu, Zhihao; Xu, Dexin; Ma, Kai; Tao, Liufeng; Xie, Zhong; Qiu, Qinjun; Pan, Shengyong (October 2022). "GeoBERTSegmenter: Word Segmentation of Chinese Texts in the Geoscience Domain Using the Improved BERT Model". Earth and Space Science. 9 (10) e2022EA002511. Bibcode:2022E&SS....902511W. doi:10.1029/2022EA002511.
^Bi, Zhen; Zhang, Ningyu; Xue, Yida; Ou, Yixin; Ji, Daxiong; Zheng, Guozhou; Chen, Huajun (2023). "OceanGPT: A Large Language Model for Ocean Science Tasks". arXiv:2310.02031 [cs.CL].
^Lin, Zhouhan; Deng, Cheng; Zhou, Le; Zhang, Tianhang; Xu, Yi; Xu, Yutong; He, Zhongmou; Shi, Yuanyuan; Dai, Beiya (2023). "GeoGalactica: A Scientific Large Language Model in Geoscience". arXiv:2401.00434 [cs.CL].
^Roberts, Jonathan; Lüddecke, Timo; Das, Sowmen; Han, Kai; Albanie, Samuel (2023). "GPT4GEO: How a Language Model Sees the World's Geography". arXiv:2306.00020 [cs.CL].
^Ji, Yuhan; Gao, Song (2023). "Evaluating the Effectiveness of Large Language Models in Representing Textual Descriptions of Geometry and Spatial Relations". arXiv:2307.03678 [cs.CL].
^Mooney, Peter; Cui, Wencong; Guan, Boyuan; Juhász, Levente (2023). "Towards Understanding the Geospatial Skills of ChatGPT: Taking a Geographic Information Systems (GIS) Exam". Proceedings of the 6th ACM SIGSPATIAL International Workshop on AI for Geographic Knowledge Discovery. pp. 85–94. doi:10.1145/3615886.3627745. ISBN979-8-4007-0348-5.
^Zhang, Yifan; Wei, Cheng; Wu, Shangyou; He, Zhengting; Yu, Wenhao (2023). "GeoGPT: Understanding and Processing Geospatial Tasks through An Autonomous GPT". arXiv:2307.07930 [cs.CL].
^Zhang, Yifan; Wang, Zhiyun; He, Zhengting; Li, Jingxuan; Mai, Gengchen; Lin, Jianfeng; Wei, Cheng; Yu, Wenhao (September 2024). "BB-GeoGPT: A framework for learning a large language model for geographic information science". Information Processing & Management. 61 (5) 103808. doi:10.1016/j.ipm.2024.103808.
^Manvi, Rohin; Khanna, Samar; Mai, Gengchen; Burke, Marshall; Lobell, David; Ermon, Stefano (2023). "GeoLLM: Extracting Geospatial Knowledge from Large Language Models". arXiv:2310.06213 [cs.CL].
^Singh, Simranjit; Fore, Michael; Stamoulis, Dimitrios (2024). "GeoLLM-Engine: A Realistic Environment for Building Geospatial Copilots". 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). pp. 585–594. doi:10.1109/CVPRW63382.2024.00063. ISBN979-8-3503-6547-4.
^Yao, Fanglong; Lu, Wanxuan; Yang, Heming; Xu, Liangyu; Liu, Chenglong; Hu, Leiyi; Yu, Hongfeng; Liu, Nayu; Deng, Chubo; Tang, Deke; Chen, Changshuo; Yu, Jiaqi; Sun, Xian; Fu, Kun (2023). "RingMo-Sense: Remote Sensing Foundation Model for Spatiotemporal Prediction via Spatiotemporal Evolution Disentangling". IEEE Transactions on Geoscience and Remote Sensing. 61: 1–21. Bibcode:2023ITGRS..6116166Y. doi:10.1109/TGRS.2023.3316166.
^Nguyen, Tung; Brandstetter, Johannes; Kapoor, Ashish; Gupta, Jayesh K.; Grover, Aditya (2023). "ClimaX: A foundation model for weather and climate". arXiv:2301.10343 [cs.LG].
^Man, Xin; Zhang, Chenghong; Feng, Jin; Li, Changyu; Shao, Jie (2023). "W-MAE: Pre-trained weather model with masked autoencoder for multi-variable weather forecasting". arXiv:2304.08754 [cs.LG].
^Guo, Xin; Lao, Jiangwei; Dang, Bo; Zhang, Yingying; Yu, Lei; Ru, Lixiang; Zhong, Liheng; Huang, Ziyuan; Wu, Kang (2023). "SkySense: A Multi-Modal Remote Sensing Foundation Model Towards Universal Interpretation for Earth Observation Imagery". arXiv:2312.10115 [cs.CV].
^Kuckreja, Kartik; Danish, Muhammad Sohail; Naseer, Muzammal; Das, Abhijit; Khan, Salman; Khan, Fahad Shahbaz (2023). "GeoChat: Grounded Large Vision-Language Model for Remote Sensing". arXiv:2311.15826 [cs.CV].
^Zhang, Wei; Cai, Miaoxin; Zhang, Tong; Zhuang, Yin; Mao, Xuerui (2024). "EarthGPT: A Universal Multimodal Large Language Model for Multisensor Image Comprehension in Remote Sensing Domain". IEEE Transactions on Geoscience and Remote Sensing. 62: 1–20. arXiv:2401.16822. Bibcode:2024ITGRS..6209624Z. doi:10.1109/TGRS.2024.3409624.
^Zhan, Yang; Xiong, Zhitong; Yuan, Yuan (March 2025). "SkyEyeGPT: Unifying remote sensing vision-language tasks via instruction tuning with large language model". ISPRS Journal of Photogrammetry and Remote Sensing. 221: 64–77. arXiv:2401.09712. Bibcode:2025JPRS..221...64Z. doi:10.1016/j.isprsjprs.2025.01.020.
^Liu, Fan; Chen, Delong; Guan, Zhangqingyun; Zhou, Xiaocong; Zhu, Jiale; Ye, Qiaolin; Fu, Liyong; Zhou, Jun (2024). "RemoteCLIP: A Vision Language Foundation Model for Remote Sensing". IEEE Transactions on Geoscience and Remote Sensing. 62: 1–16. arXiv:2306.11029. Bibcode:2024ITGRS..6290838L. doi:10.1109/TGRS.2024.3390838.
^Yuan, Zhiqiang; Zhang, Wenkai; Tian, Changyuan; Rong, Xuee; Zhang, Zhengyuan; Wang, Hongqi; Fu, Kun; Sun, Xian (2022). "Remote Sensing Cross-Modal Text-Image Retrieval Based on Global and Local Information". IEEE Transactions on Geoscience and Remote Sensing. 60: 1–16. arXiv:2204.09860. Bibcode:2022ITGRS..6063706Y. doi:10.1109/TGRS.2022.3163706.
^Hackel, Leonard; Clasen, Kai Norman; Ravanbakhsh, Mahdyar; Demir, Begüm (2023). "LIT-4-RSVQA: Lightweight Transformer-Based Visual Question Answering in Remote Sensing". IGARSS 2023 - 2023 IEEE International Geoscience and Remote Sensing Symposium. pp. 2231–2234. Bibcode:2023igar.conf..561H. doi:10.1109/IGARSS52108.2023.10281674. ISBN979-8-3503-2010-7.
^Chappuis, Christel; Zermatten, Valérie; Lobry, Sylvain; Le Saux, Bertrand; Tuia, Devis (2022). "Prompt–RSVQA: Prompting visual context to a language model for Remote Sensing Visual Question Answering". 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). pp. 1371–1380. Bibcode:2022cvpr.conf..143C. doi:10.1109/CVPRW56347.2022.00143. ISBN978-1-6654-8739-9.
^Chappuis, Christel; Mendez, Vincent; Walt, Eliot; Lobry, Sylvain; Le Saux, Bertrand; Tuia, Devis (2022). "Language Transformers for Remote Sensing Visual Question Answering". IGARSS 2022 - 2022 IEEE International Geoscience and Remote Sensing Symposium. pp. 4855–4858. Bibcode:2022igar.conf.1189C. doi:10.1109/IGARSS46834.2022.9884036. ISBN978-1-6654-2792-0.
^Chappuis, Christel; Walt, Eliot; Mendez, Vincent; Lobry, Sylvain; Saux, Bertrand Le; Tuia, Devis (2023). "The curse of language biases in remote sensing VQA: the role of spatial attributes, language diversity, and the need for clear evaluation". arXiv:2311.16782 [cs.CV].
^Xi, Zhiheng; Chen, Wenxiang; Guo, Xin; He, Wei; Ding, Yiwen; Hong, Boyang; Zhang, Ming; Wang, Junzhe; Jin, Senjie (2023). "The Rise and Potential of Large Language Model Based Agents: A Survey". arXiv:2309.07864 [cs.AI].
^Wang, Lei; Ma, Chen; Feng, Xueyang; Zhang, Zeyu; Yang, Hao; Zhang, Jingsen; Chen, Zhiyuan; Tang, Jiakai; Chen, Xu; Lin, Yankai; Zhao, Wayne Xin; Wei, Zhewei; Wen, Jirong (December 2024). "A survey on large language model based autonomous agents". Frontiers of Computer Science. 18 (6) 186345. doi:10.1007/s11704-024-40231-1.
^Guo, Haonan; Su, Xin; Wu, Chen; Du, Bo; Zhang, Liangpei; Li, Deren (2024). "Remote Sensing ChatGPT: Solving Remote Sensing Tasks with ChatGPT and Visual Models". arXiv:2401.09083 [cs.CV].
^Mitsui, Takahito; Boers, Niklas (July 2021). "Seasonal prediction of Indian summer monsoon onset with echo state networks". Environmental Research Letters. 16 (7): 074024. Bibcode:2021ERL....16g4024M. doi:10.1088/1748-9326/ac0acb.
^Lawley, Christopher J. M.; Gadd, Michael G.; Parsa, Mohammad; Lederer, Graham W.; Graham, Garth E.; Ford, Arianne (2023). "Applications of Natural Language Processing to Geoscience Text Data and Prospectivity Modeling". Natural Resources Research. 32 (4): 1503–1527. Bibcode:2023NRR....32.1503L. doi:10.1007/s11053-023-10216-1.
^Sonnewald, Maike; Lguensat, Redouane; Jones, Daniel C; Dueben, Peter D; Brajard, Julien; Balaji, V (July 2021). "Bridging observations, theory and numerical simulation of the ocean using machine learning". Environmental Research Letters. 16 (7): 073008. arXiv:2104.12506. Bibcode:2021ERL....16g3008S. doi:10.1088/1748-9326/ac0eb0.
^Ma, Chunyan; Li, Xin; Li, Yujie; Tian, Xinliang; Wang, Yichuan; Kim, Hyoungseop; Serikawa, Seiichi (2021). "Visual information processing for deep-sea visual monitoring system". Cognitive Robotics. 1: 3–11. doi:10.1016/j.cogr.2020.12.002.
Content Disclaimer
Informasi ini disarikan dari Wikipedia dan disajikan kembali untuk tujuan edukasi. Konten tersedia di bawah lisensi CC BY-SA 3.0. Kami tidak bertanggung jawab atas ketidakakuratan data yang bersumber dari kontribusi publik tersebut.
The information displayed on this website is sourced in part or in whole from Wikipedia and has been adapted for the purpose of restating it. We strive to provide accurate and relevant information, however:
There is no guarantee of absolute accuracy. Wikipedia is an open, collaborative project that can be edited by anyone, so information is subject to change.
It is not intended to constitute professional advice. The content displayed is for informational and educational purposes only. For important decisions (e.g., medical, legal, or financial), please consult a professional.
Content copyright. Wikipedia is licensed under the Creative Commons Attribution-ShareAlike License (CC BY-SA). This means that content may be reused with appropriate attribution and shared under a similar license.
Responsible use. Any risk arising from the use of information from this website is entirely the responsibility of the user.